3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
β1,056Jul 19, 2026Updated this week
Alternatives and similar repositories for MTPLX
Users that are interested in MTPLX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lossless DFlash speculative decoding for MLX on Apple Siliconβ752Jun 11, 2026Updated last month
- π₯ The fastest local AI engine for Apple Silicon. Optimised for agentic use.β82May 24, 2026Updated last month
- LLM inference server with continuous batching & SSD caching for Apple Silicon β managed from the macOS menu barβ17,950Updated this week
- Exact speculative decoding on Apple Silicon, powered by MLX.β380Apr 20, 2026Updated 3 months ago
- MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)β908Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- vLLM Metal plugin powered by mlx-swift β high-performance LLM inference on Apple Siliconβ275Jun 3, 2026Updated last month
- The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cacβ¦β3,289Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acceβ¦β62Apr 18, 2026Updated 3 months ago
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.β5,175Updated this week
- vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Bβ¦β774Updated this week
- β‘ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cacheβ¦β722May 19, 2026Updated 2 months ago
- OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batβ¦β1,446Jun 28, 2026Updated 3 weeks ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wiβ¦β146Apr 15, 2026Updated 3 months ago
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning β natively on MLX. Unslotβ¦β1,363Jun 23, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Train Large Language Models on MLX.β401Updated this week
- Free, open-source fan control for Apple Silicon Macs (M1, M2, M3, M4, M5). Menu bar app + CLI. Alternative to Macs Fan Control, TG Pro, Aβ¦β55Apr 19, 2026Updated 3 months ago
- Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agentβ¦β347Updated this week
- A native Mac App for LLM fine-tuning on Apple Silicon β fully on-device, fully open source.β246Updated this week
- JANG β GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Siliconβ213Updated this week
- DFlash: Block Diffusion for Flash Speculative Decodingβ5,496May 10, 2026Updated 2 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCmβ18,842Jul 3, 2026Updated 2 weeks ago
- Run LLMs with MLXβ6,335Jul 11, 2026Updated last week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLXβ58Mar 16, 2026Updated 4 months ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- MLX-Video is the best package for inference and finetuning of Image-Video-Audio generation models on your Mac using MLX.β273May 13, 2026Updated 2 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.β37Jun 12, 2026Updated last month
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built β¦β7,234Updated this week
- Community model zoo for Apple Core AI (iOS/macOS 27): 49 models β LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting β convertedβ¦β353Updated this week
- Fast LLM speculative inference server for consumer hardware.β2,666Updated this week
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speecβ¦β7,577Jul 10, 2026Updated last week
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon β llama.cpp forkβ118Updated this week
- Test LLMs on real tasks. Compare models side-by-side.β375Jun 16, 2026Updated last month
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAMβ785Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Local LLM Testing & Benchmarking for Apple Siliconβ197Jun 18, 2026Updated last month
- LLMs and VLMs with MLX Swiftβ68May 15, 2026Updated 2 months ago
- W8A8/W4A8 inference + optimized SDPA on Apple Silicon β unlocking unused INT8 TensorOps in M5 for 1.2β1.9Γ faster LLM prefill, plus Flashβ¦β331Jun 5, 2026Updated last month
- A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPIβ¦β352Jul 13, 2026Updated last week
- Community maintained hardware plugin for vLLM on Apple Siliconβ1,486Updated this week
- Running a big model on a small laptopβ3,988Mar 19, 2026Updated 4 months ago
- Agent Skill to help convert transformer LLMs to mlx-lmβ49Jun 16, 2026Updated last month