A repository aimed at pruning DeepSeek V3, R1 and R1-zero to a usable size
☆87Sep 5, 2025Updated 11 months ago
Alternatives and similar repositories for moe-pruner
Users that are interested in moe-pruner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official PyTorch implementation of CD-MOE☆12Mar 18, 2026Updated 5 months ago
- [ICML25] Agentic Compression Benchmark (ACBench)☆19Jul 2, 2025Updated last year
- The official implementation of the paper "Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques (TMLR)".☆89Feb 28, 2026Updated 5 months ago
- Lego for GRPO☆30May 27, 2025Updated last year
- ☆41Apr 30, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 10 months ago
- ☆17Jan 1, 2025Updated last year
- Repo hosting codes and materials related to speeding LLMs' inference using token merging.☆37Oct 9, 2025Updated 10 months ago
- Mini Model Daemon☆13Nov 9, 2024Updated last year
- Flash-Muon: An Efficient Implementation of Muon Optimizer☆260Jun 15, 2025Updated last year
- Ongoing research training transformer language models at scale, including: BERT & GPT-2☆19Jul 20, 2023Updated 3 years ago
- A FastAPI server that turns markdown prompt files into API endpoints with minimal configuration.☆15Sep 5, 2025Updated 11 months ago
- Code for "Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs"☆19Nov 6, 2025Updated 9 months ago
- The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.☆14Mar 30, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆17Nov 23, 2023Updated 2 years ago
- RWKV centralised docs for the community☆35Jan 17, 2026Updated 7 months ago
- ☆13Mar 23, 2025Updated last year
- Official Chinese documentation for RWKV | RWKV官方中文文档☆15Jun 10, 2026Updated 2 months ago
- Inference RWKV v7 in pure C.☆47Oct 10, 2025Updated 10 months ago
- A source-to-source compiler for optimizing CUDA dynamic parallelism by aggregating launches☆15Jun 21, 2019Updated 7 years ago
- ☆13Jun 29, 2024Updated 2 years ago
- RWKV v5,v6 LoRA Trainer on Cuda and Rocm Platform. RWKV is a RNN with transformer-level LLM performance. It can be directly trained like …☆13Mar 24, 2024Updated 2 years ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- LLM RP TUI for Power Users.☆36Jan 13, 2026Updated 7 months ago
- An Open Source Toolkit For LLM Distillation☆1,020May 12, 2026Updated 3 months ago
- [ACL 2024] Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models☆123May 24, 2024Updated 2 years ago
- [NAACL 2025] Representing Rule-based Chatbots with Transformers☆23Feb 9, 2025Updated last year
- [WWW 2026 Oral] MoE-CL:Self-Evolving LLMs via Continual Instruction Tuning☆21Dec 1, 2025Updated 8 months ago
- PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices☆41Jan 31, 2024Updated 2 years ago
- Love 2d based Tetris game☆13May 21, 2013Updated 13 years ago
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs☆25Nov 11, 2025Updated 9 months ago
- (ICLR 2026) Unveiling Super Experts in Mixture-of-Experts Large Language Models☆44Sep 25, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Demonstration of a factory pattern where the types automatically register themselves☆13Mar 13, 2019Updated 7 years ago
- Evolutionary strategies finetuning library for LLMs☆25Jun 29, 2026Updated last month
- Official code for the paper "Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark"☆31Jun 30, 2025Updated last year
- Official Pytorch Implementation of "Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity"☆82Jul 7, 2025Updated last year
- ☆29Aug 27, 2025Updated 11 months ago
- ☆88Jun 20, 2025Updated last year
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 10 months ago