Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Zig backend, Swift frontend macOS app with chat, music, voice, video generation.
☆1,707Oct 2, 2026Updated this week
Alternatives and similar repositories for mlx-serve
Users that are interested in mlx-serve are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B…☆2,500Sep 26, 2026Updated last week
- High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal mode…☆1,607Updated this week
- vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlm☆882Updated this week
- MLX Studio - Easiest way to run LLM's on your Mac. All in one engine.☆978Updated this week
- Open-source LLM observability and evaluation platform☆16Jul 13, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Community maintained hardware plugin for vLLM on Apple Silicon☆1,804Updated this week
- Run LLMs with MLX☆7,199Updated this week
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,561Updated this week
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆787Aug 20, 2026Updated last month
- ☆15Feb 3, 2026Updated 8 months ago
- Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server and Mac app for Apple Silicon, built on ML…☆3,874Updated this week
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆777Updated this week
- LLMs and VLMs with MLX Swift☆828Updated this week
- PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal ac…☆323Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built …☆8,015Updated this week
- Optimized LLM server configuration for Mac Studio and other Apple Silicon Macs. Headless setup with automatic startup, resource optimizat…☆359Updated this week
- A little file for doing LLM-assisted prompt expansion and image generation using Flux.schnell - complete with prompt history, prompt queu…☆26Aug 16, 2024Updated 2 years ago
- Video Frame data structures, originally part of rav1e☆18Aug 2, 2026Updated 2 months ago
- Velocity-Vortex—a fast and efficient algorithmic trading engine for the financial markets! This project is built using C++, showcasing a…☆13Apr 19, 2026Updated 5 months ago
- A high-performance, zero-copy Linear Algebra and DSP library for Apple Silicon.☆53Apr 19, 2026Updated 5 months ago
- Store, query, and create YAML workflow playbooks for LLM agents via MCP. STDIO or Streamable HTTP.☆31Sep 24, 2026Updated last week
- MLX: An array framework for Apple silicon☆64Updated this week
- Dynamically controllable Llama-model LLM inference in macOS with MLX☆18Feb 8, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Multi-role finite state machine☆31Jun 5, 2026Updated 3 months ago
- Mixed-vendor GPU inference cluster manager with speculative decoding☆37Jul 2, 2026Updated 3 months ago
- Vesta macOS Distribution - Official releases and downloads.Vesta AI Chat Assistant for macOS - Built with SwiftUI, Swift MLX and Apple I…☆82May 9, 2026Updated 4 months ago
- 🔥 The fastest local AI engine for Apple Silicon. Optimised for agentic use.☆98May 24, 2026Updated 4 months ago
- A small embeddable Datalog engine in Zig ⚡☆26Aug 1, 2026Updated 2 months ago
- MLX-Embeddings is the best package for running Vision and Language Embedding models locally on your Mac using MLX.☆449May 13, 2026Updated 4 months ago
- 'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac …☆344Sep 22, 2026Updated last week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆22,863Sep 20, 2026Updated last week
- A MCP server allowing LLM agents to easily connect and retrieve data from any database☆99Aug 1, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆224Mar 24, 2026Updated 6 months ago
- Big models. Small Macs. Zero excuses.☆61Sep 7, 2026Updated 3 weeks ago
- High-performance MLX-based LLM inference engine for macOS with native Swift implementation☆592Updated this week
- Demo project for YOLO implementation with Rust and WASM☆11Mar 29, 2024Updated 2 years ago
- Superseded by github.com/poisonxa16/pxa — PXA, the set-and-forget engine for Pascal and Volta☆24Sep 8, 2026Updated 3 weeks ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆15Aug 16, 2025Updated last year
- A cosy home for your LLMs.☆1,530Updated this week