TL;DR
- Rapid-MLX is a new open-source inference server for Apple Silicon, released under the Apache 2.0 license.
- It speaks both OpenAI’s and Anthropic’s API formats, so existing coding agents can point at it without rewriting client code.
- The project claims decode speeds up to 4x faster than Apple’s own mlx-lm, with a 1.5x gain on more typical workloads using identical model weights.
- It’s built specifically around reliable tool calling, the part of local inference that tends to break first once you hook up an autonomous coding agent.
A New Server Built for Agents, Not Just Chatbots
Rapid-MLX showed up on GitHub this week as exactly the kind of project the local-LLM crowd has been asking for without quite saying so out loud. It’s an inference server, meaning it sits between your model weights and whatever app wants to talk to them, and it’s built on top of Apple’s MLX framework specifically for Apple Silicon chips.
The project describes itself plainly: “Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents.” That last clause is doing a lot of work. This isn’t pitched as a general chatbot backend. It’s pitched at developers running coding agents locally who are tired of flaky function calls.
Two API surfaces come out of the box: OpenAI’s /v1/chat/completions and Anthropic’s /v1/messages. That means a coding agent built against either provider’s SDK can point at a local Rapid-MLX instance instead, no client-side rewrite required. And because it’s Apache 2.0, nothing stops a company from forking it, bundling it, or shipping it inside a commercial product tomorrow.
What This Means for the Local Inference Fight
Here’s the number that matters most: up to 4x faster decode speed than mlx-lm, dropping to 1.5x on what the project calls a typical task, using the same model and the same weights. That gap between the headline number and the real-world number is worth sitting with for a second. A 4x claim gets you clicks. A 1.5x claim is what you actually feel when you’re waiting on a coding agent to finish a tool call at 11pm. Big difference.
Think of mlx-lm as a stock car engine, something Apple built to run reliably and do the job without drama. Rapid-MLX is the aftermarket tuning shop that stripped it down, swapped the fuel line, and squeezed more out of the same block. It didn’t reinvent the engine. It just refused to accept the factory tuning as the ceiling.
That framing matters because of who else is in the garage. mlx-lm is Apple’s own reference implementation, the thing most Mac-based local inference defaults to. Ollama and LM Studio have spent years winning the “easiest to install” crowd, prioritizing a friendly interface over raw throughput or strict API fidelity. oMLX occupies a similar niche to Rapid-MLX itself. None of them, based on what’s public, have made tool-calling reliability for coding agents their stated headline feature the way Rapid-MLX has.
Does beating Apple’s own number mean much if nobody outside the project has reproduced it yet? Not on its own. I’ve watched enough local-LLM projects promise speed and quietly drop the “reliable tool calling” part once real agent workloads hit them, so the specificity here is what actually caught my attention, not the 4x figure. Benchmarks are easy to publish and hard to reproduce on someone else’s Mac, with someone else’s model quant, under someone else’s thermal conditions. If Rapid-MLX’s numbers hold up once more developers run it on M2 and M3 machines with their own agent setups, that’s a real win for anyone trying to cut cloud API bills on coding assistants. If they don’t, it joins a long list of local inference tools that looked great in a README and mediocre in production.
Why Apple Silicon Became the Local LLM Battleground
None of this exists without MLX, the machine learning framework Apple built to let models run efficiently on the unified memory architecture inside M-series chips. MLX turned Macs, especially the higher-memory configurations, into genuinely viable machines for running large language models without a discrete GPU or a cloud bill.
But mlx-lm is the default on-ramp, Apple’s own reference server for serving MLX models. It works, but it was never built around matching third-party API contracts exactly or squeezing out every last token per second, because that wasn’t really the job it was designed for. That left a gap for tools focused on developer workflows rather than research convenience, and coding agents are precisely the workload that exposes the gap fastest.
Coding agents call tools constantly: read this file, run this test, grep this directory, apply this patch. If the inference layer mishandles a function call format or lags badly on decode, the agent either breaks or crawls. That’s a different failure mode than a chatbot giving a slightly slow answer. If a stock server was never built with agentic tool calls in mind, why would anyone be surprised it struggles under them? It’s why a project narrowing its focus to “reliable tool calling for coding agents” rather than general chat performance is a meaningfully different bet than most local inference tools make.
The first thing worth tracking is independent benchmarking. Rapid-MLX’s 4x and 1.5x numbers come from the project itself, run on its own hardware and its own test setup, so the real test is whether other developers can reproduce anything close to those figures on their own M-series machines with their own models.
The second is adoption inside existing coding agent frameworks. Compatibility with OpenAI and Anthropic API shapes only matters if the agent tooling people already use, think editor plugins and autonomous coding assistants, actually points at it without friction. Watch for integration guides, community forks, or mentions inside popular agent repos over the next few months.
The third is whether Apple or the broader MLX maintainer community responds at all. Open-source projects that outperform a platform vendor’s own reference implementation sometimes get quietly absorbed, sometimes get ignored, and sometimes get outcompeted once the original project catches up. I’d rather see ten independent developers post real numbers than read one more README full of impressive charts, and which of those paths Rapid-MLX takes will say a lot about how seriously Apple treats the coding-agent use case on its own hardware.
Editor's Note
I have tried more local inference servers than I would like to admit, and most of them chase flashy benchmark numbers while tool calling quietly falls apart under real agent workloads. What I like here is the honesty of listing 1.5x right next to the showier 4x figure instead of burying it in a footnote. I am watching whether anyone outside the project can reproduce those numbers on their own Mac before I would trust it for anything I actually ship.
– Sanket Chaukiyal, founder, SmartChunks
FAQ
What is Rapid-MLX and who built it?
Rapid-MLX is an open-source LLM inference server for Apple Silicon Macs, built on Apple's MLX framework and released under the Apache 2.0 license. It is hosted on GitHub under the account raullenchai.
How is Rapid-MLX different from Apple's own mlx-lm?
Rapid-MLX claims decode speeds up to 4x faster than mlx-lm on some tasks and about 1.5x faster on more typical workloads, using the same model weights. It also adds OpenAI- and Anthropic-compatible API endpoints and focuses specifically on reliable tool calling for coding agents.
Does Rapid-MLX work with existing coding agents?
Yes. Because it exposes both OpenAI's /v1/chat/completions endpoint and Anthropic's /v1/messages endpoint, coding agents built against either provider's API can connect to a local Rapid-MLX instance without changing their client code.
Is Rapid-MLX free to use commercially?
Yes. It is released under the Apache 2.0 license, which permits commercial use, modification, and redistribution, including inside closed-source products.
Source: GitHub

Technology editor • 12+ years in editorial
Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.