Running a big model on a small laptop
☆3,994Mar 19, 2026Updated 4 months ago
Alternatives and similar repositories for flash-moe
Users that are interested in flash-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆122Jun 13, 2026Updated last month
- ☆218Mar 24, 2026Updated 4 months ago
- LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar☆18,222Updated this week
- ☆7,004Jul 20, 2026Updated last week
- Hundreds of models & providers. One command to find what runs on your hardware.☆30,892Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- AI agents running research on single-GPU nanochat training automatically☆92,323Mar 26, 2026Updated 4 months ago
- Training neural networks on Apple Neural Engine via reverse-engineered private APIs☆7,143Mar 10, 2026Updated 4 months ago
- Running a big model on a small laptop☆55Mar 28, 2026Updated 4 months ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆19,409Updated this week
- Run models too big for your Mac's memory☆664Apr 8, 2026Updated 3 months ago
- an open source, extensible AI agent that goes beyond code suggestions - install, execute, edit, and test with any LLM☆51,826Updated this week
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆757Jun 11, 2026Updated last month
- MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.☆5,267Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,547May 10, 2026Updated 2 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- AirLLM 70B inference with single 4GB GPU☆24,259Updated this week
- Official inference framework for 1-bit LLMs☆39,789Updated this week
- Open-source AI coworker, with memory☆16,877Updated this week
- Unsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.☆69,060Updated this week
- mac code — Claude Code, but it runs on your Mac for free. 35B AI agent at 30 tok/s via Apple Silicon flash-paging. $0/month.☆1,022Apr 9, 2026Updated 3 months ago
- Autonomous experiment loop extension for pi☆7,282Jul 15, 2026Updated 2 weeks ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆727May 19, 2026Updated 2 months ago
- Run frontier AI locally.☆46,517Jun 23, 2026Updated last month
- Fast LLM speculative inference server for consumer hardware.☆2,694Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦☆20,872Updated this week
- Adaptive Test-time Learning and Autonomous Specialization☆2,067Updated this week
- Open-source Agent Operating System☆18,062Jul 2, 2026Updated 3 weeks ago
- Foundation model for tiny devices; 14mb, 26m params, 1-6k toks/sec on mobiles, wearables smart home and robots.☆3,296Updated this week
- Fully automatic censorship removal for language models☆26,907Updated this week
- Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unslot…☆1,372Jun 23, 2026Updated last month
- The best-benchmarked open-source AI memory system. And it's free.☆57,859Updated this week
- The open-source app everyone uses to manage agents at work☆75,065Updated this week
- GLM-OCR: Accurate × Fast × Comprehensive☆7,222Apr 21, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.☆20,713Updated this week
- An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, s…☆78,149Updated this week
- A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and oth…☆30,390Updated this week
- Open-Source Frontier Voice AI☆51,268Updated this week
- Run LLMs with MLX☆6,445Updated this week
- TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecast…☆27,147Jul 14, 2026Updated 2 weeks ago
- A vector index built on TurboQuant, written in Rust with Python bindings☆14,494Updated this week