Model-agnostic MoE compression automation: build calibration bundles, run REAP/quantization/benchmark/publish stages, and render auditable reports.
☆178Mar 22, 2026Updated 4 months ago
Alternatives and similar repositories for moe-compress
Users that are interested in moe-compress are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16May 10, 2026Updated 2 months ago
- 1bit llama.cpp gguf weights paired with turboquant 4 bit kv cache☆23Apr 4, 2026Updated 4 months ago
- A pixel art space shooter built entirely by a 9B AI model on a single RTX 3060. Zero hand-written code.☆96Mar 19, 2026Updated 4 months ago
- Rust ML stack☆19Jun 27, 2026Updated last month
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Ship your repo + live coding-agent session (Claude Code / Codex / pi / Droid) to another machine over Tailscale; it resumes in tmux and k…☆170Jul 4, 2026Updated last month
- Doomscroll your way to knowledge. An infinite-scroll feed of core engineering concepts, styled like X/Twitter.☆25Apr 20, 2026Updated 3 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆845Jan 21, 2026Updated 6 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,526Updated this week
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆37Jun 12, 2026Updated last month
- Your AI friend right in your browser☆541Apr 24, 2026Updated 3 months ago
- ☆179Mar 30, 2026Updated 4 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 3 months ago
- A simple library for generating instruction tuning datasets locally☆92Jun 10, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Talk to your Obsidian vault with local models. 8ms semantic lookup via Enzyme, ~2,400 lines of TypeScript, any OpenAI-compatible endpoint…☆18Apr 3, 2026Updated 4 months ago
- A modified Codex CLI application built to support local LLMs☆72Nov 24, 2025Updated 8 months ago
- experiment loop for ai agents and swarms☆175Mar 29, 2026Updated 4 months ago
- A benchmarking harness for coding agents.☆16Updated this week
- ☆129Mar 7, 2026Updated 4 months ago
- ☆12Aug 15, 2024Updated last year
- Test LLMs on real tasks. Compare models side-by-side.☆393Jun 16, 2026Updated last month
- ☆341May 15, 2026Updated 2 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,562May 10, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Limit is your AI coding companion that refuses to leave the terminal — and that's exactly where it belongs.☆19Mar 27, 2026Updated 4 months ago
- 🥊 sparkit: The CLI & Playground for SPAR — Run AI debates in your terminal or browser☆18Jul 6, 2026Updated 3 weeks ago
- Autonomously optimize travel plans using the autoresearch pattern☆26Apr 2, 2026Updated 4 months ago
- An efficient DSA revision app that enhances learning through repetitive recall of questions.☆12Jun 30, 2023Updated 3 years ago
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆17Jul 11, 2026Updated 3 weeks ago
- Running a big model on a small laptop☆4,039Mar 19, 2026Updated 4 months ago
- ☆23May 6, 2026Updated 2 months ago
- Browse the world in the comfort of your terminal☆165Jan 8, 2026Updated 6 months ago
- ☆21Apr 6, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- Async runtime for Rust where correctness is structural: region-owned tasks, cancel-correct protocols, capability-gated effects, and deter…☆251Updated this week
- A recursive coding agent inpired by RLMs☆380Jun 22, 2026Updated last month
- A cognitive memory system for AI agents. Single SQLite file. MCP server included.☆77Jun 29, 2026Updated last month
- OpenClaw installer for Linux☆105Feb 7, 2026Updated 5 months ago
- NebulaFlow is for visually designing and running developer workflows as node graphs (CLI, LLM, control‑flow, previews). Build and execute…☆26Jul 19, 2026Updated 2 weeks ago
- Method for Long Context RLMs using verifiable Lambda Calculus☆305Apr 24, 2026Updated 3 months ago