REAP: Router-weighted Expert Activation Pruning for SMoE compression
☆490Apr 17, 2026Updated 4 months ago
Alternatives and similar repositories for reap
Users that are interested in reap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,201Updated this week
- REAM: Merging Improves Pruning of Experts in LLMs☆25Apr 16, 2026Updated 4 months ago
- A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support a…☆1,594Updated this week
- Code for data-aware compression of DeepSeek models☆76Dec 11, 2025Updated 8 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆155Aug 21, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆63Jul 10, 2025Updated last year
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,743Updated this week
- An Open Source Toolkit For LLM Distillation☆1,050May 12, 2026Updated 3 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆338Aug 19, 2026Updated last week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,816Updated this week
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,248Updated this week
- easy exllama interface w/ automation & evals☆17Aug 11, 2026Updated 2 weeks ago
- Enhancing LLMs with LoRA☆224Oct 20, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,612Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆783Updated this week
- ☆186Jul 29, 2026Updated last month
- Large-scale LLM inference engine☆1,844Aug 13, 2026Updated 2 weeks ago
- Code for the papers: “Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling” and “Adaptive Block-Scaled Data Types”☆202Apr 21, 2026Updated 4 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆460Aug 17, 2026Updated last week
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 9 months ago
- ☆77Jun 3, 2026Updated 2 months ago
- A codebase for pretraining multi-billion-scale sparse GPTs.☆28Feb 9, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆200Updated this week
- Transplants vocabulary between language models, enabling the creation of draft models for speculative decoding WITHOUT retraining.☆54Oct 29, 2025Updated 10 months ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 10 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,011Aug 18, 2026Updated last week
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆21Aug 17, 2026Updated 2 weeks ago
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 10 months ago
- Tools for merging pretrained large language models.☆7,324Jun 17, 2026Updated 2 months ago
- Official implementation for DenseMixer: Improving MoE Post-Training with Precise Router Gradient☆68Aug 3, 2025Updated last year
- DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference☆655Nov 24, 2025Updated 9 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,745Updated this week
- How much experts do we need to serve a model?☆150Mar 18, 2026Updated 5 months ago
- Fused Qwen3 MoE layer for faster training, compatible with Transformers, LoRA, bnb 4-bit quant, Unsloth. Also possible to train LoRA over…☆259Jul 24, 2026Updated last month
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 5 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,520Updated this week
- ☆29Dec 31, 2025Updated 7 months ago
- Multi-turn dataset management tool for LLM trainers☆13Mar 31, 2025Updated last year