The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
β2,559Oct 10, 2026Updated this week
Alternatives and similar repositories for MTPLX
Users that are interested in MTPLX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lossless DFlash speculative decoding for MLX on Apple Siliconβ787Aug 20, 2026Updated last month
- π₯ The fastest local AI engine for Apple Silicon. Optimised for agentic use.β99May 24, 2026Updated 4 months ago
- LLM inference server with continuous batching & SSD caching for Apple Silicon β managed from the macOS menu barβ22,687Updated this week
- Exact speculative decoding on Apple Silicon, powered by MLX.β389Apr 20, 2026Updated 5 months ago
- MLX Studio - Easiest way to run LLM's on your Mac. All in one engine.β985Updated this week
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- vLLM Metal plugin powered by mlx-swift β high-performance LLM inference on Apple Siliconβ274Jun 3, 2026Updated 4 months ago
- Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused oβ¦β3,956Updated this week
- Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Zig backend, Swift frontend macOS app with cβ¦β1,827Updated this week
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.β5,595Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acceβ¦β61Apr 18, 2026Updated 5 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wiβ¦β145Apr 15, 2026Updated 5 months ago
- High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal modeβ¦β1,618Updated this week
- vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlmβ888Updated this week
- β‘ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cacheβ¦β778Updated this week
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning β natively on MLX. Unslotβ¦β1,420Jun 23, 2026Updated 3 months ago
- Train Large Language Models on MLX.β428Updated this week
- DFlash: Block Diffusion for Flash Speculative Decodingβ6,150Aug 18, 2026Updated last month
- A native Mac App for LLM fine-tuning on Apple Silicon β fully on-device, fully open source.β274Aug 26, 2026Updated last month
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCmβ23,756Updated this week
- JANG β GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Siliconβ229Updated this week
- Run LLMs with MLXβ7,263Updated this week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLXβ58Mar 16, 2026Updated 6 months ago
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built β¦β8,031Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- MLX-Video is the best package for inference and finetuning of Image-Video-Audio generation models on your Mac using MLX.β306May 13, 2026Updated 4 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.β38Jun 12, 2026Updated 3 months ago
- Downloadable models and conversion recipes for Apple's Core AI on iPhone and Mac. Chat, vision, speech and generative models with per-modβ¦β456Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUsβ2,907Updated this week
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speecβ¦β8,022Updated this week
- Apple MLX native implementations of state-of-the-art generative image & video modelsβ2,467Updated this week
- Local LLM Testing & Benchmarking for Apple Siliconβ210Sep 5, 2026Updated last month
- A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPIβ¦β362Updated this week
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon β llama.cpp forkβ134Sep 4, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Test LLMs on real tasks. Compare models side-by-side.β429Aug 10, 2026Updated 2 months ago
- MLX Omni Server is a local inference server powered by Apple's MLX framework, specifically designed for Apple Silicon (M-series) chips. Iβ¦β745May 9, 2026Updated 5 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAMβ1,150Updated this week
- Running a big model on a small laptopβ4,236Mar 19, 2026Updated 6 months ago
- The safest, simplest way to manage Hermes from your Mac. Pure SSH. No gateways, no exposed ports, no browser layer.β2,023Jun 19, 2026Updated 3 months ago
- Community maintained hardware plugin for vLLM on Apple Siliconβ1,834Updated this week
- W8A8/W4A8 inference + optimized SDPA on Apple Silicon β unlocking unused INT8 TensorOps in M5 for 1.2β1.9Γ faster LLM prefill, plus Flashβ¦β357Jun 5, 2026Updated 4 months ago