[ICLR 2025] Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
☆163Jul 9, 2025Updated last year
Alternatives and similar repositories for DynMoE
Users that are interested in DynMoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- Inference Code for Paper "Harder Tasks Need More Experts: Dynamic Routing in MoE Models"☆74Jul 30, 2024Updated 2 years ago
- ☆20May 28, 2025Updated last year
- [NeurIPS 2024] Efficiency for Free: Ideal Data Are Transportable Representations☆19Jan 19, 2025Updated last year
- [ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts☆280Oct 16, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [Preprint] GMem: A Modular Approach for Ultra-Efficient Generative Models☆43Mar 11, 2025Updated last year
- [ACL 2026 Main] Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis☆52Jun 30, 2026Updated 2 months ago
- [ICLR 2026] Any-step Generation via N-th Order Recursive Consistent Velocity Field Estimation☆42Feb 4, 2026Updated 7 months ago
- Official implementation of RMoE (Layerwise Recurrent Router for Mixture-of-Experts)☆33Aug 4, 2024Updated 2 years ago
- [ICML 2023] FedBR: Improving Federated Learning on Heterogeneous Data via Local Learning Bias Reduction☆29Mar 7, 2024Updated 2 years ago
- [ ICLR 2025 ] Making LLMs More Effective with Hierarchical Mixture of LoRA Experts☆32Oct 9, 2025Updated 11 months ago
- Community Implementation of the paper: "Multi-Head Mixture-of-Experts" In PyTorch☆31Aug 29, 2026Updated 3 weeks ago
- The official implementation of the paper "Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques (TMLR)".☆91Aug 22, 2026Updated 3 weeks ago
- [CVPR 2026] IOMM: Fast Pre-training of Unified Multimodal Models without Text-Image Pairs☆26Apr 11, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs☆25Nov 11, 2025Updated 10 months ago
- [TKDE'25] The official GitHub page for the survey paper "A Survey on Mixture of Experts in Large Language Models".☆511Aug 18, 2026Updated last month
- ☆17Jun 11, 2025Updated last year
- The code of Advancing Expert Specialization for Better MoE (NeurIPS2025 oral)☆34Jan 22, 2026Updated 7 months ago
- Official code for the ICLR 2025 paper, "Ada-K Routing: Boosting the Efficiency of MoE-based LLMs"☆12Mar 1, 2025Updated last year
- [ICLR 2023] "Sparse MoE as the New Dropout: Scaling Dense and Self-Slimmable Transformers" by Tianlong Chen*, Zhenyu Zhang*, Ajay Jaiswal…☆56Feb 28, 2023Updated 3 years ago
- PyTorch implementation of "From Sparse to Soft Mixtures of Experts"☆72Aug 22, 2023Updated 3 years ago
- ☆146Jul 21, 2024Updated 2 years ago
- ☆25Oct 22, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The source code of "Merging Experts into One: Improving Computational Efficiency of Mixture of Experts (EMNLP 2023)":☆47Aug 22, 2026Updated 3 weeks ago
- A comprehensive framework for benchmarking single and multi-agent systems across a wide range of tasks—evaluating performance, accuracy, …☆38Jun 5, 2026Updated 3 months ago
- This repository contains the code for the paper "ST-MoE-BERT: A Spatial-Temporal Mixture-of-Experts Framework for Long-Term Cross-City Mo…☆16Feb 20, 2025Updated last year
- [NeurIPS 2024] Search for Efficient LLMs☆16Jan 16, 2025Updated last year
- [ICLR 2025] Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization☆25Oct 5, 2025Updated 11 months ago
- [ICLR2025] γ -MOD: Mixture-of-Depth Adaptation for Multimodal Large Language Models☆46Oct 28, 2025Updated 10 months ago
- The codebase for our EMNLP24 paper: Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Mo…☆85Jan 27, 2025Updated last year
- Implementation of the "the first large-scale multimodal mixture of experts models." from the paper: "Multimodal Contrastive Learning with…☆36Aug 29, 2026Updated 3 weeks ago
- A collection of AWESOME things about mixture-of-experts☆1,288Dec 8, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆23Nov 26, 2024Updated last year
- Official PyTorch Implementation of EMoE: Unlocking Emergent Modularity in Large Language Models [main conference @ NAACL2024]☆40May 28, 2024Updated 2 years ago
- Code release for AdapMoE accepted by ICCAD 2024☆39Apr 28, 2025Updated last year
- [ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.☆123Dec 20, 2024Updated last year
- ☆140Jun 6, 2025Updated last year
- ☆39Jan 16, 2025Updated last year
- Survey: A collection of AWESOME papers and resources on the latest research in Mixture of Experts.☆146Aug 21, 2024Updated 2 years ago