The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
β2,380Sep 19, 2026Updated this week
Alternatives and similar repositories for MTPLX
Users that are interested in MTPLX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lossless DFlash speculative decoding for MLX on Apple Siliconβ785Aug 20, 2026Updated last month
- π₯ The fastest local AI engine for Apple Silicon. Optimised for agentic use.β98May 24, 2026Updated 3 months ago
- LLM inference server with continuous batching & SSD caching for Apple Silicon β managed from the macOS menu barβ21,914Updated this week
- Exact speculative decoding on Apple Silicon, powered by MLX.β390Apr 20, 2026Updated 5 months ago
- MLX Studio - Easiest way to run LLM's on your Mac. All in one engine.β973Updated this week
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- vLLM Metal plugin powered by mlx-swift β high-performance LLM inference on Apple Siliconβ275Jun 3, 2026Updated 3 months ago
- The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cacβ¦β3,791Updated this week
- Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agentβ¦β1,405Updated this week
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.β5,511Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acceβ¦β61Apr 18, 2026Updated 5 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wiβ¦β145Apr 15, 2026Updated 5 months ago
- High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal modeβ¦β1,587Updated this week
- vMLX - Use MLX models easily - JANGQ (GGUF for MLX) - Not dependant on mlx_vlmβ861Updated this week
- β‘ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cacheβ¦β768Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning β natively on MLX. Unslotβ¦β1,412Jun 23, 2026Updated 2 months ago
- Train Large Language Models on MLX.β414Updated this week
- DFlash: Block Diffusion for Flash Speculative Decodingβ6,102Aug 18, 2026Updated last month
- A native Mac App for LLM fine-tuning on Apple Silicon β fully on-device, fully open source.β267Aug 26, 2026Updated 3 weeks ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCmβ22,519Updated this week
- JANG β GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Siliconβ227Sep 4, 2026Updated 2 weeks ago
- Run LLMs with MLXβ7,065Updated this week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLXβ58Mar 16, 2026Updated 6 months ago
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built β¦β7,971Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- MLX-Video is the best package for inference and finetuning of Image-Video-Audio generation models on your Mac using MLX.β301May 13, 2026Updated 4 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.β38Jun 12, 2026Updated 3 months ago
- Downloadable models and conversion recipes for Apple's Core AI on iPhone and Mac. Chat, vision, speech and generative models with per-modβ¦β438Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUsβ2,868Updated this week
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speecβ¦β7,914Updated this week
- Apple MLX native implementations of state-of-the-art generative image & video modelsβ2,335Updated this week
- Local LLM Testing & Benchmarking for Apple Siliconβ207Sep 5, 2026Updated 2 weeks ago
- A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPIβ¦β361Updated this week
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon β llama.cpp forkβ134Sep 4, 2026Updated 2 weeks ago
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Test LLMs on real tasks. Compare models side-by-side.β421Aug 10, 2026Updated last month
- MLX Omni Server is a local inference server powered by Apple's MLX framework, specifically designed for Apple Silicon (M-series) chips. Iβ¦β746May 9, 2026Updated 4 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAMβ1,096Sep 13, 2026Updated last week
- Running a big model on a small laptopβ4,144Mar 19, 2026Updated 6 months ago
- The safest, simplest way to manage Hermes from your Mac. Pure SSH. No gateways, no exposed ports, no browser layer.β2,020Jun 19, 2026Updated 3 months ago
- Community maintained hardware plugin for vLLM on Apple Siliconβ1,747Updated this week
- W8A8/W4A8 inference + optimized SDPA on Apple Silicon β unlocking unused INT8 TensorOps in M5 for 1.2β1.9Γ faster LLM prefill, plus Flashβ¦β348Jun 5, 2026Updated 3 months ago