REAP: Router-weighted Expert Activation Pruning for SMoE compression
☆447Apr 17, 2026Updated 3 months ago
Alternatives and similar repositories for reap
Users that are interested in reap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,056Updated this week
- REAM: Merging Improves Pruning of Experts in LLMs☆22Apr 16, 2026Updated 3 months ago
- A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support…☆1,532Updated this week
- Code for data-aware compression of DeepSeek models☆75Dec 11, 2025Updated 7 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆146Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆63Jul 10, 2025Updated last year
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,566Updated this week
- An Open Source Toolkit For LLM Distillation☆986May 12, 2026Updated 2 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆2,943Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆326Jul 1, 2026Updated 2 weeks ago
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,211Updated this week
- easy exllama interface w/ automation & evals☆17Updated this week
- Enhancing LLMs with LoRA☆224Oct 20, 2025Updated 9 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,278Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆633Updated this week
- ☆161Updated this week
- Large-scale LLM inference engine☆1,806Updated this week
- Code for the papers: “Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling” and “Adaptive Block-Scaled Data Types”☆198Apr 21, 2026Updated 3 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆397May 29, 2026Updated last month
- Find the hidden meaning of LLMs☆41Nov 13, 2025Updated 8 months ago
- ☆68Jun 3, 2026Updated last month
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆191Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Transplants vocabulary between language models, enabling the creation of draft models for speculative decoding WITHOUT retraining.☆54Oct 29, 2025Updated 8 months ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 8 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,504May 10, 2026Updated 2 months ago
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆17Jul 5, 2026Updated 2 weeks ago
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 9 months ago
- Tools for merging pretrained large language models.☆7,250Jun 17, 2026Updated last month
- Official implementation for DenseMixer: Improving MoE Post-Training with Precise Router Gradient☆67Aug 3, 2025Updated 11 months ago
- DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference☆648Nov 24, 2025Updated 7 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,481Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- How much experts do we need to serve a model?☆153Mar 18, 2026Updated 4 months ago
- Fused Qwen3 MoE layer for faster training, compatible with Transformers, LoRA, bnb 4-bit quant, Unsloth. Also possible to train LoRA over…☆260Updated this week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 4 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,087Updated this week
- ☆28Dec 31, 2025Updated 6 months ago
- Multi-turn dataset management tool for LLM trainers☆13Mar 31, 2025Updated last year
- A PyTorch framework for training transformer language models with Mixture of Experts (MoE) architecture support, Mixture of Depths (MoD),…☆21Updated this week