Model-agnostic MoE compression automation: build calibration bundles, run REAP/quantization/benchmark/publish stages, and render auditable reports.
☆181Mar 22, 2026Updated 5 months ago
Alternatives and similar repositories for moe-compress
Users that are interested in moe-compress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17May 10, 2026Updated 4 months ago
- 1bit llama.cpp gguf weights paired with turboquant 4 bit kv cache☆24Apr 4, 2026Updated 5 months ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆99Mar 19, 2026Updated 5 months ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 4 months ago
- Ship your repo + live coding-agent session (Claude Code / Codex / pi / Droid) to another machine over Tailscale; it resumes in tmux and k…☆171Sep 3, 2026Updated last week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Doomscroll your way to knowledge. An infinite-scroll feed of core engineering concepts, styled like X/Twitter.☆25Apr 20, 2026Updated 4 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆1,279Sep 3, 2026Updated last week
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,764Sep 2, 2026Updated last week
- LocalAGI agent hub☆16Mar 16, 2026Updated 5 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 3 months ago
- ☆181Mar 30, 2026Updated 5 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 4 months ago
- A simple library for generating instruction tuning datasets locally☆110Aug 21, 2026Updated 3 weeks ago
- Talk to your Obsidian vault with local models. 8ms semantic lookup via Enzyme, ~2,400 lines of TypeScript, any OpenAI-compatible endpoint…☆18Apr 3, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A modified Codex CLI application built to support local LLMs☆73Nov 24, 2025Updated 9 months ago
- experiment loop for ai agents and swarms☆174Mar 29, 2026Updated 5 months ago
- A benchmarking harness for coding agents.☆16Aug 25, 2026Updated 2 weeks ago
- ☆130Mar 7, 2026Updated 6 months ago
- Xilly Game Mode is a competitive-grade optimization utility designed to instantly reallocate your PC's resources for maximum gaming perfo…☆20Feb 13, 2026Updated 7 months ago
- ☆29Mar 14, 2026Updated 5 months ago
- ☆342May 15, 2026Updated 3 months ago
- ☆16Updated this week
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,076Aug 18, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Limit is your AI coding companion that refuses to leave the terminal — and that's exactly where it belongs.☆19Mar 27, 2026Updated 5 months ago
- 🥊 sparkit: The CLI & Playground for SPAR — Run AI debates in your terminal or browser☆18Jul 6, 2026Updated 2 months ago
- ArtAgents☆16Updated this week
- A tiny Rust gateway for running coding agents across model providers safely.☆75Aug 30, 2026Updated 2 weeks ago
- Running a big model on a small laptop☆4,093Mar 19, 2026Updated 5 months ago
- ☆23May 6, 2026Updated 4 months ago
- Browse the world in the comfort of your terminal☆165Jan 8, 2026Updated 8 months ago
- ☆12Jul 8, 2024Updated 2 years ago
- NanoClaw (Venice API) — Personal AI assistant powered by Venice AI. Fork of NanoClaw with WhatsApp + Telegram support.☆19Mar 7, 2026Updated 6 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,851Updated this week
- A recursive coding agent inpired by RLMs☆389Jun 22, 2026Updated 2 months ago
- A cognitive memory system for AI agents. Single SQLite file. MCP server included.☆79Aug 6, 2026Updated last month
- Anthropic-compatible HTTP facade over claude-agent-acp☆92Jun 20, 2026Updated 2 months ago
- Add Speech to Text to your Omarchy (Arch Linux) System☆29Sep 17, 2025Updated 11 months ago
- OpenClaw installer for Linux☆104Feb 7, 2026Updated 7 months ago
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆49Jul 11, 2026Updated 2 months ago