Model-agnostic MoE compression automation: build calibration bundles, run REAP/quantization/benchmark/publish stages, and render auditable reports.
☆181Mar 22, 2026Updated 6 months ago
Alternatives and similar repositories for moe-compress
Users that are interested in moe-compress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 1bit llama.cpp gguf weights paired with turboquant 4 bit kv cache☆24Apr 4, 2026Updated 6 months ago
- Fast Opportunistic Mixture-Of-Experts. From-scratch C/HIP MoE inference with multi-tier caching and cache-aware routing. First ever examp…☆31Mar 31, 2026Updated 6 months ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆106Mar 19, 2026Updated 6 months ago
- Rust ML stack☆19Sep 14, 2026Updated 3 weeks ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Ship your repo + live coding-agent session (Claude Code / Codex / pi / Droid) to another machine over Tailscale; it resumes in tmux and k…☆175Sep 3, 2026Updated last month
- Doomscroll your way to knowledge. An infinite-scroll feed of core engineering concepts, styled like X/Twitter.☆25Apr 20, 2026Updated 5 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆1,345Updated this week
- Codebase for the paper "Biases in the Blind Spot: Detecting What LLMs Fail to Mention"☆19May 28, 2026Updated 4 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,822Updated this week
- LocalAGI agent hub☆16Mar 16, 2026Updated 6 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 3 months ago
- Dynamic Telegram Trading Bot☆20Feb 21, 2025Updated last year
- Your AI friend right in your browser☆550Sep 17, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆183Mar 30, 2026Updated 6 months ago
- Talk to your Obsidian vault with local models. 8ms semantic lookup via Enzyme, ~2,400 lines of TypeScript, any OpenAI-compatible endpoint…☆19Apr 3, 2026Updated 6 months ago
- A simple library for generating instruction tuning datasets locally☆118Sep 20, 2026Updated 2 weeks ago
- Erdős Problem #870: paper and sorry-free, axiom-clean Lean 4 formalization.☆21Jun 26, 2026Updated 3 months ago
- A benchmarking harness for coding agents.☆17Sep 23, 2026Updated 2 weeks ago
- ☆130Mar 7, 2026Updated 7 months ago
- experiment loop for ai agents and swarms☆173Mar 29, 2026Updated 6 months ago
- Rust CLI for RBMEM (.rbmem), a structured Rust-Brain memory format with timestamp-protected sections, hierarchy, graph relations, Hermes …☆28Jun 14, 2026Updated 3 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆429Aug 10, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Turbo1Bit: Combining 1-bit LLM weights (Bonsai) with TurboQuant KV cache compression for maximum inference efficiency. 4.2x KV cache comp…☆31Apr 2, 2026Updated 6 months ago
- ☆340May 15, 2026Updated 4 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,144Aug 18, 2026Updated last month
- A tiny Rust gateway for running coding agents across model providers safely.☆73Aug 30, 2026Updated last month
- A ComfyUI plugin that provides a user interface of StableStudio☆23Aug 15, 2025Updated last year
- ☆23May 6, 2026Updated 5 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,904Updated this week
- Async runtime for Rust where correctness is structural: region-owned tasks, cancel-correct protocols, capability-gated effects, and deter…☆281Updated this week
- A recursive coding agent inpired by RLMs☆391Jun 22, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A cognitive memory system for AI agents. Single SQLite file. MCP server included.☆79Aug 6, 2026Updated 2 months ago
- Repo for measuring whether using AI tools inhibits skill formation and development☆16Jan 3, 2026Updated 9 months ago
- Anthropic-compatible HTTP facade over claude-agent-acp☆91Sep 19, 2026Updated 2 weeks ago
- Add Speech to Text to your Omarchy (Arch Linux) System☆28Sep 17, 2025Updated last year
- OpenClaw installer for Linux☆104Feb 7, 2026Updated 8 months ago
- Method for Long Context RLMs using verifiable Lambda Calculus☆308Apr 24, 2026Updated 5 months ago
- A thinking process, not a template. 4-stage planning framework with 21 anti-pattern checks, adversarial hardening, and 8-domain detection…☆46Aug 24, 2026Updated last month