The GPU-free LLM inference engine. Combines lazy expert loading + TurboQuant KV compression to run models that shouldn't fit on your hardware. Built from scratch, fully local, zero cloud.
☆23Apr 13, 2026Updated 5 months ago
Alternatives and similar repositories for lazy-moe
Users that are interested in lazy-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cohere Toolkit is a collection of prebuilt components enabling users to quickly build and deploy RAG applications.☆30Jan 19, 2025Updated last year
- Lightweight reverse proxy + admin UI that turns your backend endpoints into multi-tenant Model Context Protocol tools for OpenServ—or any…☆16Sep 8, 2026Updated 3 weeks ago
- ARK (Automated Resource Knowledge-base) revolutionizes personal computing by creating an open-source, decentralized assistant. Harnessing…☆19Sep 19, 2026Updated 2 weeks ago
- A Ray Tracing-Inspired Approach to Neural Network Optimization☆17Jun 11, 2025Updated last year
- The mono repo of tools needed to convert MCP servers into Claude Code Skills☆15Nov 25, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Exception Details - Capture, and Inspect and all variables from Exception-time.☆22Aug 11, 2013Updated 13 years ago
- Python Data Controller for Neural EEG headsets. (Windows + Linux)☆12Dec 27, 2017Updated 8 years ago
- ☆27Jun 11, 2025Updated last year
- Official Zitadel auth example for Flutter.☆14May 6, 2026Updated 4 months ago
- A deterministic object hashing algorithm for Node.js.