3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
β1,811Aug 30, 2026Updated this week
Alternatives and similar repositories for MTPLX
Users that are interested in MTPLX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Lossless DFlash speculative decoding for MLX on Apple Siliconβ776Aug 20, 2026Updated last week
- π₯ The fastest local AI engine for Apple Silicon. Optimised for agentic use.β94May 24, 2026Updated 3 months ago
- LLM inference server with continuous batching & SSD caching for Apple Silicon β managed from the macOS menu barβ20,999Updated this week
- Exact speculative decoding on Apple Silicon, powered by MLX.β388Apr 20, 2026Updated 4 months ago
- MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)β960Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- vLLM Metal plugin powered by mlx-swift β high-performance LLM inference on Apple Siliconβ275Jun 3, 2026Updated 2 months ago
- The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cacβ¦β3,590Updated this week
- Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agentβ¦β948Updated this week
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.β5,440Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acceβ¦β61Apr 18, 2026Updated 4 months ago
- vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Bβ¦β835Updated this week
- β‘ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cacheβ¦β752Updated this week
- High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal modeβ¦β1,552Updated this week
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wiβ¦β146Apr 15, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning β natively on MLX. Unslotβ¦β1,391Jun 23, 2026Updated 2 months ago
- Train Large Language Models on MLX.β409Updated this week
- Free, open-source fan control for Apple Silicon Macs (M1, M2, M3, M4, M5). Menu bar app + CLI. Alternative to Macs Fan Control, TG Pro, Aβ¦β96Updated this week
- A native Mac App for LLM fine-tuning on Apple Silicon β fully on-device, fully open source.β258Updated this week
- JANG β GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Siliconβ224Updated this week
- DFlash: Block Diffusion for Flash Speculative Decodingβ6,011Aug 18, 2026Updated last week
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCmβ21,927Updated this week
- Run LLMs with MLXβ6,840Updated this week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLXβ58Mar 16, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- MLX-Video is the best package for inference and finetuning of Image-Video-Audio generation models on your Mac using MLX.β293May 13, 2026Updated 3 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.β38Jun 12, 2026Updated 2 months ago
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built β¦β7,740Updated this week
- Community model zoo for Apple Core AI (iOS/macOS 27): 62 models β LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting β each gateβ¦β399Updated this week
- Apple MLX native implementations of state-of-the-art generative image & video modelsβ2,298Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUsβ2,816Updated this week
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speecβ¦β7,808Updated this week
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon β llama.cpp forkβ129Aug 15, 2026Updated 2 weeks ago
- Test LLMs on real tasks. Compare models side-by-side.β415Aug 10, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAMβ985Updated this week
- Local LLM Testing & Benchmarking for Apple Siliconβ203Aug 23, 2026Updated last week
- LLMs and VLMs with MLX Swiftβ66May 15, 2026Updated 3 months ago
- W8A8/W4A8 inference + optimized SDPA on Apple Silicon β unlocking unused INT8 TensorOps in M5 for 1.2β1.9Γ faster LLM prefill, plus Flashβ¦β340Jun 5, 2026Updated 2 months ago
- Community maintained hardware plugin for vLLM on Apple Siliconβ1,666Updated this week
- A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPIβ¦β358Updated this week
- Running a big model on a small laptopβ4,087Mar 19, 2026Updated 5 months ago