The GPU-free LLM inference engine. Combines lazy expert loading + TurboQuant KV compression to run models that shouldn't fit on your hardware. Built from scratch, fully local, zero cloud.
☆23Apr 13, 2026Updated 3 months ago
Alternatives and similar repositories for lazy-moe
Users that are interested in lazy-moe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting: