The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Rapid-MLX is a local AI inference server for Apple Silicon Macs that exposes an OpenAI-compatible HTTP endpoint, allowing any tool that targets the OpenAI API to run against locally hosted models instead. It installs via pip, Homebrew, or a one-line curl script, then serves models including Qwen3.5, Gemma 4, DeepSeek V4, and GPT-OSS through a FastAPI backend built on Apple's MLX framework. Claude Code users point their base URL to localhost:8000 to get fully local, zero-cost inference, and the server also works with Cursor and Aider the same way. The project ships 17 tool-call parsers covering multiple model families, a prompt cache yielding 0.08-second cached time-to-first-token, and a reasoning-separation feature that strips chain-of-thought tokens before returning responses. A cloud routing option falls back to remote providers when needed. Benchmarks on an M3 Ultra show GPT-OSS 20B running at 119 tokens per second, and the project claims 2.3x throughput over Ollama under four-concurrent-user load using identical model weights. Developers running local coding assistants on macOS hardware with 16 GB or more of unified memory are the primary audience.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Healthy fork ratio
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
git clone https://github.com/raullenchai/Rapid-MLXTools overview
What people ask about Rapid-MLX
What is raullenchai/Rapid-MLX?
+
raullenchai/Rapid-MLX is tools for the Claude AI ecosystem. The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider. It has 3.4k GitHub stars and was last updated today.
How do I install Rapid-MLX?
+
You can install Rapid-MLX by cloning the repository (https://github.com/raullenchai/Rapid-MLX) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is raullenchai/Rapid-MLX safe to use?
+
Our security agent has analyzed raullenchai/Rapid-MLX and assigned a Trust Score of 100/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains raullenchai/Rapid-MLX?
+
raullenchai/Rapid-MLX is maintained by raullenchai. The last recorded GitHub activity is from today, with 38 open issues.
Are there alternatives to Rapid-MLX?
+
Yes. On ClaudeWave you can browse similar tools at /categories/tools, sorted by popularity or recent activity.
Deploy Rapid-MLX to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/raullenchai-rapid-mlx)<a href="https://claudewave.com/repo/raullenchai-rapid-mlx"><img src="https://claudewave.com/api/badge/raullenchai-rapid-mlx" alt="Featured on ClaudeWave: raullenchai/Rapid-MLX" width="320" height="64" /></a>More Tools
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
An AI SKILL that provide design intelligence for building professional UI/UX multiple platforms
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary