REAP: Router-weighted Expert Activation Pruning for SMoE compression
☆507Apr 17, 2026Updated 5 months ago
Alternatives and similar repositories for reap
Users that are interested in reap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,459Updated this week
- REAM: Merging Improves Pruning of Experts in LLMs☆26Apr 16, 2026Updated 5 months ago
- A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包☆1,621Updated this week
- Code for data-aware compression of DeepSeek models☆75Dec 11, 2025Updated 9 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆166Sep 11, 2026Updated last week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆63Jul 10, 2025Updated last year
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,799Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,244Updated this week
- An Open Source Toolkit For LLM Distillation☆1,066May 12, 2026Updated 4 months ago
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆340Aug 19, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,868Updated this week
- easy exllama interface w/ automation & evals☆19Updated this week
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,257Updated this week
- Enhancing LLMs with LoRA☆225Oct 20, 2025Updated 11 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,828Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆843Updated this week
- Large-scale LLM inference engine☆1,861Sep 11, 2026Updated last week
- ☆194Jul 29, 2026Updated last month
- Code for the papers: “Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling” and “Adaptive Block-Scaled Data Types”☆203Apr 21, 2026Updated 4 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆471Aug 17, 2026Updated last month
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 10 months ago
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆203Sep 2, 2026Updated 2 weeks ago
- ☆86Jun 3, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Transplants vocabulary between language models, enabling the creation of draft models for speculative decoding WITHOUT retraining.☆54Oct 29, 2025Updated 10 months ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 10 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,102Aug 18, 2026Updated last month
- A codebase for pretraining multi-billion-scale sparse GPTs.☆30Feb 9, 2026Updated 7 months ago
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 11 months ago
- Tools for merging pretrained large language models.☆7,356Sep 12, 2026Updated last week
- Official implementation for DenseMixer: Improving MoE Post-Training with Precise Router Gradient☆68Aug 3, 2025Updated last year
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆22Aug 17, 2026Updated last month
- DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference☆656Nov 24, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,698Updated this week
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,768Updated this week
- How much experts do we need to serve a model?☆151Mar 18, 2026Updated 6 months ago
- Fused Qwen3 MoE layer for faster training, compatible with Transformers, LoRA, bnb 4-bit quant, Unsloth. Also possible to train LoRA over…☆259Jul 24, 2026Updated last month
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 6 months ago
- ☆29Dec 31, 2025Updated 8 months ago
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆21Updated this week