☆222Mar 24, 2026Updated 5 months ago
Alternatives and similar repositories for flash-moe
Users that are interested in flash-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 4 months ago
- Running a big model on a small laptop☆57Mar 28, 2026Updated 5 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 5 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆131Updated this week
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆131Jun 13, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Minimal Claude Code alternative powered by MLX☆47Jan 11, 2026Updated 7 months ago
- Flash-MoE iOS — Run massive MoE models on iPhone☆49Mar 23, 2026Updated 5 months ago
- Running a big model on a small laptop☆4,087Mar 19, 2026Updated 5 months ago
- Run transformers directly on Apple's Neural Engine in Swift bypass coreml entirely☆167Aug 24, 2026Updated 2 weeks ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 2 months ago
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated last year
- ☆32May 15, 2026Updated 3 months ago
- ☆86Mar 3, 2026Updated 6 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆778Aug 20, 2026Updated 2 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- This repo maintains a 'cheat sheet' for LLMs that are undertrained on mlx☆33Mar 12, 2026Updated 5 months ago
- ☆45Mar 5, 2026Updated 6 months ago
- Run models too big for your Mac's memory☆666Aug 27, 2026Updated last week
- Local AI runtime for training & running small LLMs directly on Apple Neural Engine (ANE). No CoreML. No Metal. Offline, on-device fine-tu…☆121Aug 23, 2026Updated 2 weeks ago
- Find out why your CoreML model isn't running on the Neural Engine!☆30Jun 18, 2024Updated 2 years ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆16Aug 16, 2025Updated last year
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆31Updated this week
- Community model zoo for Apple Core AI (iOS/macOS 27): 66 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gate…☆409Updated this week
- Artificial Neural Engine Machine Learning Library☆1,656Mar 10, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,401Jun 23, 2026Updated 2 months ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆760Updated this week
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated last year
- CLI to demonstrate running a large language model (LLM) on Apple Neural Engine.☆131Dec 27, 2024Updated last year
- MLX Model Manager unifies loading and inferencing with LLMs and VLMs.☆101Jan 30, 2025Updated last year
- Rust-native hybrid training & inference engine for Apple Neural Engine + Metal GPU☆181Apr 3, 2026Updated 5 months ago
- Perplexica is an AI-powered search engine. It is an Open source alternative to Perplexity AI☆18Feb 22, 2026Updated 6 months ago
- `Paper repo for “Coherence-Guided Dead-Head Identification in Frozen Transformers,” including manuscript sources, figures, frozen result …☆51Apr 8, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Train Large Language Models on MLX.☆411Updated this week
- ☆18May 27, 2025Updated last year
- 3.34× faster inference on Apple Silicon — native MLX port of DFlash speculative decoding☆19Apr 11, 2026Updated 4 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 4 months ago
- A minimalistic Swift implementation of the Jinja templating engine, specifically designed for parsing and rendering ML chat templates.☆131Jul 21, 2026Updated last month
- LLM training on Apple's Neural Engine — native Obj-C, private APIs, zero GPU. Dynamic weight pipeline for training without kernel recompi…☆58Mar 17, 2026Updated 5 months ago
- ☆21Oct 9, 2024Updated last year