Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
☆418Aug 3, 2026Updated this week
Alternatives and similar repositories for mlx-serve
Users that are interested in mlx-serve are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.☆1,125Updated this week
- vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont B…☆785Updated this week
- OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous bat…☆1,485Updated this week
- LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar☆18,407Updated this week
- Open-source LLM observability and evaluation platform☆15Jul 13, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)☆924Updated this week
- wasm-interface-types supplement & compiler of wasmedge☆18Aug 25, 2023Updated 2 years ago
- Community maintained hardware plugin for vLLM on Apple Silicon☆1,540Updated this week
- Run LLMs with MLX☆6,509Jul 26, 2026Updated last week
- LLMs and VLMs with MLX Swift☆764Updated this week
- The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cac…☆3,392Updated this week
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆756Jun 11, 2026Updated last month
- Optimized Ollama LLM server configuration for Mac Studio and other Apple Silicon Macs. Headless setup with automatic startup, resource op…☆347Jan 24, 2026Updated 6 months ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆730May 19, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A little file for doing LLM-assisted prompt expansion and image generation using Flux.schnell - complete with prompt history, prompt queu…☆26Aug 16, 2024Updated last year
- Video Frame data structures, originally part of rav1e☆18Jun 21, 2026Updated last month
- Velocity-Vortex—a fast and efficient algorithmic trading engine for the financial markets! This project is built using C++, showcasing a…☆13Apr 19, 2026Updated 3 months ago
- A hands-on Swift and Metal course for building LLM inference from first principles on Apple silicon, with 48 guided lessons, runnable exe…☆182Jul 17, 2026Updated 2 weeks ago
- A high-performance, zero-copy Linear Algebra and DSP library for Apple Silicon.☆53Apr 19, 2026Updated 3 months ago
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built …☆7,484Updated this week
- Beautiful GUI for Postgres optimized for macOS☆16Jan 14, 2026Updated 6 months ago
- Store, query, and create YAML workflow playbooks for LLM agents via MCP. STDIO or Streamable HTTP.☆31Updated this week
- MLX: An array framework for Apple silicon☆53Jul 23, 2026Updated last week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Multi-role finite state machine☆29Jun 5, 2026Updated last month
- Vesta macOS Distribution - Official releases and downloads.Vesta AI Chat Assistant for macOS - Built with SwiftUI, Swift MLX and Apple I…☆84May 9, 2026Updated 2 months ago
- A small embeddable Datalog engine in Zig ⚡☆24Updated this week
- MLX-Embeddings is the best package for running Vision and Language Embedding models locally on your Mac using MLX.☆423May 13, 2026Updated 2 months ago
- A MCP server allowing LLM agents to easily connect and retrieve data from any database☆99Aug 1, 2025Updated last year
- Big models. Small Macs. Zero excuses.☆59Updated this week
- Run LLMs with MLX☆15May 1, 2026Updated 3 months ago
- Demo project for YOLO implementation with Rust and WASM☆11Mar 29, 2024Updated 2 years ago
- A WebSocket implementation based on Swift Concurrency☆11Sep 13, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac …☆323Updated this week
- Simple and opinionated web framework written in zig☆21Jun 3, 2026Updated 2 months ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆16Aug 16, 2025Updated 11 months ago
- High-performance MLX-based LLM inference engine for macOS with native Swift implementation☆576Jun 8, 2026Updated last month
- A Deribit Trading System using C++ utilizing WebSocket that performs the following actions: Place an order Cancel an order Modify an ord…☆12Oct 2, 2025Updated 10 months ago
- Tool calling using MLX Swift across iOS, macOS, and visionOS platforms☆125Jul 2, 2026Updated last month
- Utilities to evaluate MLX quantizations☆22Jun 9, 2026Updated last month