REAP: Router-weighted Expert Activation Pruning for SMoE compression
☆525Apr 17, 2026Updated 5 months ago
Alternatives and similar repositories for reap
Users that are interested in reap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,624Updated this week
- REAM: Merging Improves Pruning of Experts in LLMs☆27Apr 16, 2026Updated 5 months ago
- A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包☆1,630Updated this week
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆170Sep 11, 2026Updated 3 weeks ago
- ☆64Jul 10, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- State-of-the-art LLM compression, built for production inference with vLLM☆3,859Updated this week
- An Open Source Toolkit For LLM Distillation☆1,079May 12, 2026Updated 4 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,288Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆342Aug 19, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,904Updated this week
- easy exllama interface w/ automation & evals☆19Sep 25, 2026Updated 2 weeks ago
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,270Updated this week
- Enhancing LLMs with LoRA☆228Oct 20, 2025Updated 11 months ago
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆5,251Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆865Updated this week
- Large-scale LLM inference engine☆1,870Sep 11, 2026Updated 3 weeks ago
- ☆199Jul 29, 2026Updated 2 months ago
- Code for the papers: “Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling” and “Adaptive Block-Scaled Data Types”☆204Apr 21, 2026Updated 5 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆476Aug 17, 2026Updated last month
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 10 months ago
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆203Sep 2, 2026Updated last month
- ☆93Jun 3, 2026Updated 4 months ago
- Transplants vocabulary between language models, enabling the creation of draft models for speculative decoding WITHOUT retraining.☆54Oct 29, 2025Updated 11 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 11 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,144Aug 18, 2026Updated last month
- A codebase for pretraining multi-billion-scale sparse GPTs.☆30Feb 9, 2026Updated 8 months ago
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 11 months ago
- Tools for merging pretrained large language models.☆7,392Sep 12, 2026Updated 3 weeks ago
- Official implementation for DenseMixer: Improving MoE Post-Training with Precise Router Gradient☆68Aug 3, 2025Updated last year
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆24Aug 17, 2026Updated last month
- DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference☆661Nov 24, 2025Updated 10 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,906Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,822Updated this week
- How much experts do we need to serve a model?☆152Mar 18, 2026Updated 6 months ago
- Fused Qwen3 MoE layer for faster training, compatible with Transformers, LoRA, bnb 4-bit quant, Unsloth. Also possible to train LoRA over…☆259Jul 24, 2026Updated 2 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 6 months ago
- ☆29Dec 31, 2025Updated 9 months ago
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆20Updated this week
- A PyTorch implementation of gradient-free optimization for directly optimizing NDCG (Normalized Discounted Cumulative Gain) in neural inf…☆19Dec 21, 2025Updated 9 months ago