Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Up to 4× faster than Apple's MLX (mlx-lm) on the same weights.
☆3,941Oct 8, 2026Updated this week
Alternatives and similar repositories for Rapid-MLX
Users that are interested in Rapid-MLX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B…☆2,542Updated this week
- LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar☆22,637Updated this week
- High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal mode…☆1,616Updated this week
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆787Aug 20, 2026Updated last month
- 🔥 The fastest local AI engine for Apple Silicon. Optimised for agentic use.☆98May 24, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- MLX Studio - Easiest way to run LLM's on your Mac. All in one engine.☆984Updated this week
- Exact speculative decoding on Apple Silicon, powered by MLX.☆389Apr 20, 2026Updated 5 months ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆777Updated this week
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built …☆8,027Updated this week
- vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm☆887Updated this week
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,417Jun 23, 2026Updated 3 months ago
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,588Updated this week
- Run LLMs with MLX☆7,252Updated this week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆23,698Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,143Aug 18, 2026Updated last month
- ⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.☆5,952Aug 30, 2026Updated last month
- The headless browser for AI agents and web scraping☆28,697Updated this week
- A tool for creating and running Linux containers using lightweight virtual machines on a Mac. It is written in Swift, and optimized for A…☆50,577Updated this week
- 🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent bec…☆100,006Updated this week
- The agent that grows with you☆252,152Updated this week
- Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.☆29,041Updated this week
- Make humans and AI agents work as one team — open-source and self-hostable.☆52,195Updated this week
- Production-grade engineering skills for AI coding agents.☆103,398Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🐹 Clean, uninstall, analyze, optimize, and monitor your Mac. Free open-source CLI, plus a native Mac app.☆69,647Updated this week
- 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: http…☆86,342Updated this week
- A slide framework built for agents.☆9,087Updated this week
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speec…☆8,013Updated this week
- #1 Persistent memory for AI coding agents based on real-world benchmarks☆29,242Updated this week
- PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.☆29,498Updated this week
- Open-Source Frontier Voice AI☆54,678Updated this week
- Desktop app to manage markdown knowledge bases☆20,002Updated this week
- Garry's Opinionated OpenClaw/Hermes Agent Brain☆30,684Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning☆38,382Sep 30, 2026Updated last week
- OpenHuman is the fastest, cheapest, most efficient open-source agent harness. Written in Rust☆41,673Updated this week
- MLX: An array framework for Apple silicon☆28,690Updated this week
- Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions abo…☆85,604Updated this week
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆274Jun 3, 2026Updated 4 months ago
- AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI☆113,482Updated this week
- Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.☆13,773Sep 9, 2026Updated last month