Running a big model on a small laptop
☆4,072Mar 19, 2026Updated 4 months ago
Alternatives and similar repositories for flash-moe
Users that are interested in flash-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆126Jun 13, 2026Updated 2 months ago
- ☆221Mar 24, 2026Updated 4 months ago
- LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar☆19,392Updated this week
- ☆7,004Jul 20, 2026Updated 3 weeks ago
- Hundreds of models & providers. One command to find what runs on your hardware.☆32,727Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- AI agents running research on single-GPU nanochat training automatically☆94,072Mar 26, 2026Updated 4 months ago
- Training neural networks on Apple Neural Engine via reverse-engineered private APIs☆7,239Mar 10, 2026Updated 5 months ago
- Running a big model on a small laptop☆56Mar 28, 2026Updated 4 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆21,528Aug 9, 2026Updated last week
- Run models too big for your Mac's memory☆666Apr 8, 2026Updated 4 months ago
- an open source, extensible AI agent that goes beyond code suggestions - install, execute, edit, and test with any LLM☆52,955Updated this week
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆766Jun 11, 2026Updated 2 months ago
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,362Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,643May 10, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- AirLLM 70B inference with single 4GB GPU☆31,581Updated this week
- Official inference framework for 1-bit LLMs☆40,102Jul 27, 2026Updated 3 weeks ago
- Open-source AI coworker, with memory☆17,316Updated this week
- Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.☆73,589Updated this week
- mac code — Claude Code, but it runs on your Mac for free. 35B AI agent at 30 tok/s via Apple Silicon flash-paging. $0/month.☆1,024Apr 9, 2026Updated 4 months ago
- Autonomous experiment loop extension for pi☆7,634Jul 15, 2026Updated last month
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆742Updated this week
- Run frontier AI locally.☆46,872Jun 23, 2026Updated last month
- LLM speculative inference server for consumer & heterogeneous hardware☆2,762Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦☆25,406Updated this week
- Adaptive Test-time Learning and Autonomous Specialization☆2,077Updated this week
- Open-source Agent Operating System☆18,120Jul 2, 2026Updated last month
- Fully automatic censorship removal for language models☆27,802Updated this week
- 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.☆7,505Updated this week
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,382Jun 23, 2026Updated last month
- The best-benchmarked open-source AI memory system. And it's free.☆58,435Updated this week
- The open-source app everyone uses to manage agents at work☆78,770Updated this week
- GLM-OCR: Accurate × Fast × Comprehensive☆7,290Apr 21, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.☆21,468Updated this week
- An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, s…☆80,254Updated this week
- A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and oth…☆30,542Updated this week
- Open-Source Frontier Voice AI☆52,873Jul 24, 2026Updated 3 weeks ago
- Run LLMs with MLX☆6,683Updated this week
- TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecast…☆28,006Jul 14, 2026Updated last month
- A vector index built on TurboQuant, written in Rust with Python bindings☆14,853Updated this week