⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.
☆777Oct 1, 2026Updated this week
Alternatives and similar repositories for SwiftLM
Users that are interested in SwiftLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆275Jun 3, 2026Updated 4 months ago
- Exact speculative decoding on Apple Silicon, powered by MLX.☆391Apr 20, 2026Updated 5 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆787Aug 20, 2026Updated last month
- ☆38Mar 30, 2026Updated 6 months ago
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B…☆2,500Sep 26, 2026Updated last week
- LLMs and VLMs with MLX Swift☆828Updated this week
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆137Jun 13, 2026Updated 3 months ago
- LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar☆22,455Updated this week
- High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal mode…☆1,607Updated this week
- MLX Studio - Easiest way to run LLM's on your Mac. All in one engine.☆978Updated this week
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,561Updated this week
- Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers☆30May 19, 2026Updated 4 months ago
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,414Jun 23, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built …☆8,015Updated this week
- Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server and Mac app for Apple Silicon, built on ML…☆3,874Updated this week
- AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML☆1,209Updated this week
- An API-compatible, drop-in replacement for Apple's Foundation Models framework with support for custom language model providers.