3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
β1,144Aug 7, 2026Updated this week
Alternatives and similar repositories for MTPLX
Users that are interested in MTPLX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lossless DFlash speculative decoding for MLX on Apple Siliconβ758Jun 11, 2026Updated last month
- π₯ The fastest local AI engine for Apple Silicon. Optimised for agentic use.β88May 24, 2026Updated 2 months ago
- LLM inference server with continuous batching & SSD caching for Apple Silicon β managed from the macOS menu barβ18,529Updated this week
- Exact speculative decoding on Apple Silicon, powered by MLX.β382Apr 20, 2026Updated 3 months ago
- MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)β930Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- vLLM Metal plugin powered by mlx-swift β high-performance LLM inference on Apple Siliconβ274Jun 3, 2026Updated 2 months ago
- The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cacβ¦β3,419Updated this week
- Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agentβ¦β545Updated this week
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.β5,308Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acceβ¦β62Apr 18, 2026Updated 3 months ago
- vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Bβ¦β792Updated this week
- β‘ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cacheβ¦β734Updated this week
- OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batβ¦β1,497Updated this week
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wiβ¦β145Apr 15, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning β natively on MLX. Unslotβ¦β1,377Jun 23, 2026Updated last month
- Train Large Language Models on MLX.β407Jul 21, 2026Updated 2 weeks ago
- A native Mac App for LLM fine-tuning on Apple Silicon β fully on-device, fully open source.β254Jul 16, 2026Updated 3 weeks ago
- JANG β GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Siliconβ218Updated this week
- DFlash: Block Diffusion for Flash Speculative Decodingβ5,575May 10, 2026Updated 2 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCmβ20,977Updated this week
- Run LLMs with MLXβ6,544Updated this week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLXβ58Mar 16, 2026Updated 4 months ago
- MLX-Video is the best package for inference and finetuning of Image-Video-Audio generation models on your Mac using MLX.β284May 13, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.β37Jun 12, 2026Updated last month
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built β¦β7,566Updated this week
- Community model zoo for Apple Core AI (iOS/macOS 27): 57 models β LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting β each gateβ¦β383Updated this week
- MLX native implementations of state-of-the-art generative image modelsβ2,268Updated this week
- LLM speculative inference server for consumer hardware & heterogeneous computingβ2,725Updated this week
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speecβ¦β7,694Updated this week
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon β llama.cpp forkβ118Jul 15, 2026Updated 3 weeks ago
- Test LLMs on real tasks. Compare models side-by-side.β394Jun 16, 2026Updated last month
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAMβ839Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Local LLM Testing & Benchmarking for Apple Siliconβ199Jun 18, 2026Updated last month
- LLMs and VLMs with MLX Swiftβ67May 15, 2026Updated 2 months ago
- W8A8/W4A8 inference + optimized SDPA on Apple Silicon β unlocking unused INT8 TensorOps in M5 for 1.2β1.9Γ faster LLM prefill, plus Flashβ¦β337Jun 5, 2026Updated 2 months ago
- A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPIβ¦β356Updated this week
- Community maintained hardware plugin for vLLM on Apple Siliconβ1,554Updated this week
- Running a big model on a small laptopβ4,049Mar 19, 2026Updated 4 months ago
- Agent Skill to help convert transformer LLMs to mlx-lmβ51Jun 16, 2026Updated last month