Awesome Mixture of Experts (MoE): A Curated List of Mixture of Experts (MoE) and Mixture of Multimodal Experts (MoME)
☆73Jul 13, 2026Updated last month
Alternatives and similar repositories for Awesome-Mixture-of-Experts
Users that are interested in Awesome-Mixture-of-Experts are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Survey: A collection of AWESOME papers and resources on the latest research in Mixture of Experts.☆146Aug 21, 2024Updated 2 years ago
- CRAI is a multimodal large language model based on the Mixture of Experts (MoE) architecture, supporting text and image cross-modal tasks…☆16Apr 29, 2025Updated last year
- The code of Advancing Expert Specialization for Better MoE (NeurIPS2025 oral)☆35Jan 22, 2026Updated 7 months ago
- Mixture of Experts from scratch☆14Apr 12, 2024Updated 2 years ago
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- The code of 《M4: Multi-Proxy Multi-Gate Mixture of Experts Network for Multiple Instance Learning in Histopathology Image Analysis》☆14Mar 31, 2025Updated last year
- [NeurIPS 2024] Mixture of Experts for Audio-Visual Learning☆25Jan 19, 2025Updated last year
- [CVPR 2025] CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answeri…☆62Jun 16, 2025Updated last year
- Multi-Source Domain Attention☆14Oct 21, 2020Updated 5 years ago
- Scaling Laws for Mixture of Experts Models☆15Feb 25, 2025Updated last year
- A curated reading list of research in Mixture-of-Experts(MoE).☆670Oct 30, 2024Updated last year
- Community Implementation of the paper: "Multi-Head Mixture-of-Experts" In PyTorch☆31Updated this week
- [TKDE'25] The official GitHub page for the survey paper "A Survey on Mixture of Experts in Large Language Models".☆505Aug 18, 2026Updated 2 weeks ago
- [ICML 2025] Code for "R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts"☆19Mar 10, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Mixture-of-Experts Multimodal Variational Autoencoder☆15Jul 3, 2025Updated last year
- The official implementation of "MoCaE: Mixture of Calibrated Experts Significantly Improves Accuracy in Object Detection"☆47Mar 25, 2025Updated last year
- ☆19Jul 17, 2026Updated last month
- 吴恩达深度学习课程笔记PyTorch版☆17Nov 2, 2025Updated 10 months ago
- GraphAlign: Pretraining One Graph Neural Network on Multiple Graphs via Feature Alignment☆19Sep 17, 2024Updated last year
- Code for "Inducer-tuning: Connecting Prefix-tuning and Adapter-tuning" (EMNLP 2022) and "Empowering Parameter-Efficient Transfer Learning…☆11Feb 6, 2023Updated 3 years ago
- Code for the paper "No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations"☆11Oct 31, 2024Updated last year
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 3 months ago
- A curated list of medical reasoning research on large language models, organized by modality, technique, application, and benchmark.☆18Oct 17, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation of "Mixture of Experts Meets Prompt-Based Continual Learning" (NeurIPS 2024)☆43Aug 1, 2025Updated last year
- This repository contains the code for the paper "ST-MoE-BERT: A Spatial-Temporal Mixture-of-Experts Framework for Long-Term Cross-City Mo…☆16Feb 20, 2025Updated last year
- MoE-Visualizer is a tool designed to visualize the selection of experts in Mixture-of-Experts (MoE) models.☆16Apr 8, 2025Updated last year
- COSE: Configuring Serverless Functions using Statistical Learning☆10Jun 28, 2023Updated 3 years ago
- MineRL Navigate Video Dataset☆13Mar 24, 2021Updated 5 years ago
- Code for kdd-24 paper "GPFedRec: Graph-Guided Personalization for Federated Recommendation"☆22Dec 2, 2024Updated last year
- Official Pytorch implementation of 'Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning'? (ICLR2024)☆13Mar 8, 2024Updated 2 years ago
- Source codes for paper "BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity".☆19Jan 10, 2026Updated 7 months ago
- Mamba R1 represents a novel architecture that combines the efficiency of Mamba's state space models with the scalability of Mixture of Ex…☆25Oct 13, 2025Updated 10 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICML 2025] AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism☆20Jul 14, 2025Updated last year
- Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision Language Models☆24Oct 12, 2025Updated 10 months ago
- Cross Visual Prompt Tuning [ICCV 2025]☆13Aug 3, 2025Updated last year
- [CVPR 2023] Unofficial PyTorch implementation for CVPR2023 paper, Prototypical Residual Networks for Anomaly Detection and Localization.☆39Jul 7, 2023Updated 3 years ago
- 使用fastrtc框架调用qwen-2.5-omni-realtime实现实时语音、视频等☆14Jun 27, 2025Updated last year
- Sequoia: Enabling Quality-of-Service in Serverless Computing☆11Sep 14, 2020Updated 5 years ago
- Contains implementation of the DoubIL and ResiduIL algorithms from the ICML '22 paper Causal Imitation Learning under Temporally Correlat…☆11Dec 9, 2022Updated 3 years ago