☆224Mar 24, 2026Updated 6 months ago
Alternatives and similar repositories for flash-moe
Users that are interested in flash-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆38Mar 30, 2026Updated 5 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 5 months ago
- Running a big model on a small laptop☆57Mar 28, 2026Updated 5 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 6 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆134Sep 4, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆134Jun 13, 2026Updated 3 months ago
- Minimal Claude Code alternative powered by MLX☆47Jan 11, 2026Updated 8 months ago
- Running a big model on a small laptop☆4,151Mar 19, 2026Updated 6 months ago
- Flash-MoE iOS — Run massive MoE models on iPhone☆58Mar 23, 2026Updated 6 months ago
- Run transformers directly on Apple's Neural Engine in Swift bypass coreml entirely☆172Sep 12, 2026Updated 2 weeks ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 3 months ago
- ☆32May 15, 2026Updated 4 months ago
- ☆86Mar 3, 2026Updated 6 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆786Aug 20, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This repo maintains a 'cheat sheet' for LLMs that are undertrained on mlx☆33Mar 12, 2026Updated 6 months ago
- ☆45Mar 5, 2026Updated 6 months ago
- Run models too big for your Mac's memory☆668Aug 27, 2026Updated last month
- Find out why your CoreML model isn't running on the Neural Engine!☆30Jun 18, 2024Updated 2 years ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆15Aug 16, 2025Updated last year
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆31Sep 6, 2026Updated 3 weeks ago
- Artificial Neural Engine Machine Learning Library☆1,680Sep 18, 2026Updated last week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 5 months ago
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,413Jun 23, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated last year
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆774Updated this week
- Run LLMs on Apple devices with CoreML, optimized for Apple Neural Engine + GPU☆195Sep 14, 2026Updated 2 weeks ago
- A repo of useful MLX skills.☆89Jan 25, 2026Updated 8 months ago
- CLI to demonstrate running a large language model (LLM) on Apple Neural Engine.☆131Dec 27, 2024Updated last year
- Run 70B+ LLMs on Apple Silicon by using SSD as extended memory — intelligent layer streaming and caching for Mac☆49Aug 24, 2026Updated last month
- MLX Model Manager unifies loading and inferencing with LLMs and VLMs.