The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Rapid-MLX is a local AI inference server for Apple Silicon Macs that exposes an OpenAI-compatible HTTP endpoint, allowing any tool that targets the OpenAI API to run against locally hosted models instead. It installs via pip, Homebrew, or a one-line curl script, then serves models including Qwen3.5, Gemma 4, DeepSeek V4, and GPT-OSS through a FastAPI backend built on Apple's MLX framework. Claude Code users point their base URL to localhost:8000 to get fully local, zero-cost inference, and the server also works with Cursor and Aider the same way. The project ships 17 tool-call parsers covering multiple model families, a prompt cache yielding 0.08-second cached time-to-first-token, and a reasoning-separation feature that strips chain-of-thought tokens before returning responses. A cloud routing option falls back to remote providers when needed. Benchmarks on an M3 Ultra show GPT-OSS 20B running at 119 tokens per second, and the project claims 2.3x throughput over Ollama under four-concurrent-user load using identical model weights. Developers running local coding assistants on macOS hardware with 16 GB or more of unified memory are the primary audience.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Healthy fork ratio
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
git clone https://github.com/raullenchai/Rapid-MLXResumen de Tools
Lo que la gente pregunta sobre Rapid-MLX
¿Qué es raullenchai/Rapid-MLX?
+
raullenchai/Rapid-MLX es tools para el ecosistema de Claude AI. The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider. Tiene 3.4k estrellas en GitHub y se actualizó por última vez today.
¿Cómo se instala Rapid-MLX?
+
Puedes instalar Rapid-MLX clonando el repositorio (https://github.com/raullenchai/Rapid-MLX) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar raullenchai/Rapid-MLX?
+
Nuestro agente de seguridad ha analizado raullenchai/Rapid-MLX y le ha asignado un Trust Score de 100/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene raullenchai/Rapid-MLX?
+
raullenchai/Rapid-MLX es mantenido por raullenchai. La última actividad registrada en GitHub es de today, con 38 issues abiertos.
¿Hay alternativas a Rapid-MLX?
+
Sí. En ClaudeWave puedes explorar tools similares en /categories/tools, ordenados por popularidad o actividad reciente.
Despliega Rapid-MLX en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/raullenchai-rapid-mlx)<a href="https://claudewave.com/repo/raullenchai-rapid-mlx"><img src="https://claudewave.com/api/badge/raullenchai-rapid-mlx" alt="Featured on ClaudeWave: raullenchai/Rapid-MLX" width="320" height="64" /></a>Más Tools
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
An AI SKILL that provide design intelligence for building professional UI/UX multiple platforms
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary