vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
☆785Aug 3, 2026Updated this week
Alternatives and similar repositories for vmlx
Users that are interested in vmlx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)☆924Updated this week
- JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon☆218Updated this week
- LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar☆18,407Updated this week
- 3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.☆1,125Updated this week
- OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous bat…☆1,485Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆756Jun 11, 2026Updated last month
- Run LLMs with MLX☆6,509Jul 26, 2026Updated last week
- Exact speculative decoding on Apple Silicon, powered by MLX.☆382Apr 20, 2026Updated 3 months ago
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built …☆7,484Updated this week
- Community maintained hardware plugin for vLLM on Apple Silicon☆1,540Updated this week
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,291Updated this week
- The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cac…☆3,392Updated this week
- fastest runtime for apple silicon.☆89Apr 16, 2026Updated 3 months ago
- Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent…☆418Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆274Jun 3, 2026Updated 2 months ago
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,376Jun 23, 2026Updated last month
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆730May 19, 2026Updated 2 months ago
- A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speec…☆7,675Updated this week
- Apple Silicon (MLX) port of Karpathy's autoresearch — autonomous AI research loops on Mac, no PyTorch required.☆1,778Jul 2, 2026Updated last month
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 3 months ago
- MLX-Embeddings is the best package for running Vision and Language Embedding models locally on your Mac using MLX.☆423May 13, 2026Updated 2 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 3 months ago
- MLX-Video is the best package for inference and finetuning of Image-Video-Audio generation models on your Mac using MLX.☆282May 13, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPI…☆356Jul 27, 2026Updated last week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆326Jul 1, 2026Updated last month
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆20,357Updated this week
- Attempt at Porting LTX-2 Video Model to Apple's MLX Machine Learning Framework☆116Apr 18, 2026Updated 3 months ago
- MLX native implementations of state-of-the-art generative image models☆2,260Updated this week
- MLX: An array framework for Apple silicon☆27,804Updated this week
- 🔥 The fastest local AI engine for Apple Silicon. Optimised for agentic use.☆88May 24, 2026Updated 2 months ago
- Train Large Language Models on MLX.☆405Jul 21, 2026Updated last week
- ☆32May 15, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LM Studio Apple MLX engine☆1,137Updated this week
- Running a big model on a small laptop☆4,039Mar 19, 2026Updated 4 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,562May 10, 2026Updated 2 months ago
- Local LLM Testing & Benchmarking for Apple Silicon☆199Jun 18, 2026Updated last month
- Utilities to evaluate MLX quantizations☆22Jun 9, 2026Updated last month
- MLX-GUI MLX Inference Server for Apple Silicone☆209Apr 1, 2026Updated 4 months ago
- Artificial Neural Engine Machine Learning Library☆1,632Mar 10, 2026Updated 4 months ago