REAP: Router-weighted Expert Activation Pruning for SMoE compression
☆474Apr 17, 2026Updated 3 months ago
Alternatives and similar repositories for reap
Users that are interested in reap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,127Updated this week
- REAM: Merging Improves Pruning of Experts in LLMs☆22Apr 16, 2026Updated 3 months ago
- A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support…☆1,560Updated this week
- Code for data-aware compression of DeepSeek models☆76Dec 11, 2025Updated 7 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆148Aug 1, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆63Jul 10, 2025Updated last year
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,652Updated this week
- An Open Source Toolkit For LLM Distillation☆998May 12, 2026Updated 2 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,024Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆327Jul 1, 2026Updated last month
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,728Updated this week
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,223Updated this week
- easy exllama interface w/ automation & evals☆17Aug 2, 2026Updated last week
- Enhancing LLMs with LoRA☆224Oct 20, 2025Updated 9 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,418Updated this week
- A codebase for pretraining multi-billion-scale sparse GPTs.☆27Feb 9, 2026Updated 6 months ago
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆711Updated this week
- ☆169Jul 29, 2026Updated last week
- Large-scale LLM inference engine☆1,825Updated this week
- Code for the papers: “Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling” and “Adaptive Block-Scaled Data Types”☆202Apr 21, 2026Updated 3 months ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆431Jul 25, 2026Updated 2 weeks ago
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 8 months ago
- ☆73Jun 3, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆197Jul 20, 2026Updated 3 weeks ago
- Transplants vocabulary between language models, enabling the creation of draft models for speculative decoding WITHOUT retraining.☆54Oct 29, 2025Updated 9 months ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 9 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,579May 10, 2026Updated 3 months ago
- Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90%…☆18Jul 5, 2026Updated last month
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 9 months ago
- Tools for merging pretrained large language models.☆7,287Jun 17, 2026Updated last month
- DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference☆654Nov 24, 2025Updated 8 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,638Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- How much experts do we need to serve a model?☆152Mar 18, 2026Updated 4 months ago
- Fused Qwen3 MoE layer for faster training, compatible with Transformers, LoRA, bnb 4-bit quant, Unsloth. Also possible to train LoRA over…☆259Jul 24, 2026Updated 2 weeks ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 4 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,312Updated this week
- ☆28Dec 31, 2025Updated 7 months ago
- Multi-turn dataset management tool for LLM trainers☆13Mar 31, 2025Updated last year
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆67Jul 11, 2026Updated last month