☆219Mar 24, 2026Updated 4 months ago
Alternatives and similar repositories for flash-moe
Users that are interested in flash-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆36Mar 30, 2026Updated 3 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 3 months ago
- Running a big model on a small laptop☆55Mar 28, 2026Updated 4 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 4 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆117Jul 15, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆122Jun 13, 2026Updated last month
- Minimal Claude Code alternative powered by MLX☆47Jan 11, 2026Updated 6 months ago
- Flash-MoE iOS — Run massive MoE models on iPhone☆48Mar 23, 2026Updated 4 months ago
- Running a big model on a small laptop☆3,994Mar 19, 2026Updated 4 months ago
- Train and run transformers directly on Apple's Neural Engine in Swift bypass coreml entirely☆157Jul 16, 2026Updated last week
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆36Jun 12, 2026Updated last month
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated 11 months ago
- ☆31May 15, 2026Updated 2 months ago
- ☆86Mar 3, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆757Jun 11, 2026Updated last month
- This repo maintains a 'cheat sheet' for LLMs that are undertrained on mlx☆33Mar 12, 2026Updated 4 months ago
- ☆42Mar 5, 2026Updated 4 months ago
- Run models too big for your Mac's memory☆664Apr 8, 2026Updated 3 months ago
- Local AI runtime for training & running small LLMs directly on Apple Neural Engine (ANE). No CoreML. No Metal. Offline, on-device fine-tu…☆110Mar 6, 2026Updated 4 months ago
- Find out why your CoreML model isn't running on the Neural Engine!☆30Jun 18, 2024Updated 2 years ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆16Aug 16, 2025Updated 11 months ago
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆29Jul 18, 2026Updated last week
- Community model zoo for Apple Core AI (iOS/macOS 27): 57 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gate…☆362Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Multi-LoRA inference server for Apple Silicon -- one base model, many adapters, zero reload☆19Apr 13, 2026Updated 3 months ago
- Artificial Neural Engine Machine Learning Library☆1,630Mar 10, 2026Updated 4 months ago
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,375Jun 23, 2026Updated last month
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆728May 19, 2026Updated 2 months ago
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated 10 months ago
- A repo of useful MLX skills.☆87Jan 25, 2026Updated 6 months ago
- CLI to demonstrate running a large language model (LLM) on Apple Neural Engine.☆131Dec 27, 2024Updated last year
- MLX Model Manager unifies loading and inferencing with LLMs and VLMs.☆102Jan 30, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Rust-native hybrid training & inference engine for Apple Neural Engine + Metal GPU☆179Apr 3, 2026Updated 3 months ago
- Perplexica is an AI-powered search engine. It is an Open source alternative to Perplexity AI☆18Feb 22, 2026Updated 5 months ago
- `Paper repo for “Coherence-Guided Dead-Head Identification in Frozen Transformers,” including manuscript sources, figures, frozen result …☆49Apr 8, 2026Updated 3 months ago
- Train Large Language Models on MLX.☆402Jul 21, 2026Updated last week
- ☆18May 27, 2025Updated last year
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 3 months ago
- A minimalistic Swift implementation of the Jinja templating engine, specifically designed for parsing and rendering ML chat templates.☆131Jul 21, 2026Updated last week