High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Healthy fork ratio
- ✓Clear description
- ✓Topics declared
git clone https://github.com/waybarrios/vllm-mlxTools overview
What people ask about vllm-mlx
What is waybarrios/vllm-mlx?
+
waybarrios/vllm-mlx is tools for the Claude AI ecosystem. High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support. It has 1.6k GitHub stars and its last recorded update is dated 2026-09-06.
How do I install vllm-mlx?
+
You can install vllm-mlx by cloning the repository (https://github.com/waybarrios/vllm-mlx) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is waybarrios/vllm-mlx safe to use?
+
Our security agent has analyzed waybarrios/vllm-mlx and assigned a Trust Score of 97/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains waybarrios/vllm-mlx?
+
waybarrios/vllm-mlx is maintained by waybarrios. The last recorded GitHub activity is dated 2026-09-06, with 105 open issues.
Are there alternatives to vllm-mlx?
+
Yes. On ClaudeWave you can browse similar tools at /categories/tools, sorted by popularity or recent activity.
Deploy vllm-mlx to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/waybarrios-vllm-mlx)<a href="https://claudewave.com/repo/waybarrios-vllm-mlx"><img src="https://claudewave.com/api/badge/waybarrios-vllm-mlx" alt="Featured on ClaudeWave: waybarrios/vllm-mlx" width="320" height="64" /></a>More Tools
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
An AI skill that provides design intelligence for building professional UI/UX across multiple platforms.
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
Use Claude Code, Codex, Pi, and OpenCode and more for free (1.3B+ free tokens) from your terminal, app, IDE, or phone like OpenClaw (voice supported + ToS friendly)