The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
☆3,561Aug 28, 2026Updated this week
Alternatives and similar repositories for Rapid-MLX
Users that are interested in Rapid-MLX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.☆1,734Updated this week
- LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar☆20,879Updated this week
- High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal mode…☆1,550Updated this week
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆774Aug 20, 2026Updated last week
- 🔥 The fastest local AI engine for Apple Silicon. Optimised for agentic use.☆93May 24, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)☆959Updated this week
- Exact speculative decoding on Apple Silicon, powered by MLX.☆386Apr 20, 2026Updated 4 months ago
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built …☆7,699Updated this week
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆752Updated this week
- vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont B…☆832Updated this week
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,390Jun 23, 2026Updated 2 months ago
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,433Updated this week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆21,882Updated this week
- Run LLMs with MLX☆6,821Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,997Aug 18, 2026Updated last week
- ⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.☆5,810Updated this week
- The headless browser for AI agents and web scraping☆22,388Updated this week
- A tool for creating and running Linux containers using lightweight virtual machines on a Mac. It is written in Swift, and optimized for A…☆49,478Updated this week
- 🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent bec…☆92,318Updated this week
- The agent that grows with you☆237,651Updated this week
- Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.☆21,982Updated this week
- Make humans and AI agents work as one team — open-source and self-hostable.☆48,123Updated this week
- Production-grade engineering skills for AI coding agents.☆90,456Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 🐹 Clean, uninstall, analyze, optimize, and monitor your Mac. Free open-source CLI, plus a native Mac app.☆65,183Updated this week
- 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!☆76,915Updated this week
- A slide framework built for agents.☆7,238Updated this week
- Open-Source Frontier Voice AI☆53,292Jul 24, 2026Updated last month
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speec…☆7,800Updated this week
- #1 Persistent memory for AI coding agents based on real-world benchmarks☆27,667Updated this week
- PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.☆28,847Updated this week
- Desktop app to manage markdown knowledge bases☆19,608Updated this week
- Garry's Opinionated OpenClaw/Hermes Agent Brain☆29,236Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning☆36,214Updated this week
- Your Personal AI super intelligence. A brain that builds a local-first memory of your life, a fantastic orchestrator of agent fleets and …☆38,724Updated this week
- Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions abo…☆80,854Updated this week
- MLX: An array framework for Apple silicon☆28,202Updated this week
- AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI☆98,764Updated this week
- Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.☆13,735Jul 24, 2026Updated last month
- Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization,…☆26,551Updated this week