The code of Advancing Expert Specialization for Better MoE (NeurIPS2025 oral)
☆34Jan 22, 2026Updated 8 months ago
Alternatives and similar repositories for Auxloss-For-Advancing-Expert-Specialization
Users that are interested in Auxloss-For-Advancing-Expert-Specialization are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Compositional Muon release☆25Jun 5, 2026Updated 3 months ago
- Official Pytorch Implementation of "Outlier-weighed Layerwise Sampling for LLM Fine-tuning" by Pengxiang Li, Lu Yin, Xiaowei Gao, Shiwei …☆35Jun 3, 2025Updated last year
- Expert Specialization MoE Solution based on CUTLASS☆27Apr 14, 2026Updated 5 months ago
- [ICML 2025] Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts☆37Nov 10, 2025Updated 10 months ago
- ☆24Dec 6, 2025Updated 9 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 5 months ago
- Post-Trained MoE Can Skip Half Experts via Self-Distillation☆41Sep 6, 2026Updated 2 weeks ago
- Repo for the paper: GamutNet: Restoring Wide-Gamut Colors for Camera-Captured Images☆10Feb 8, 2023Updated 3 years ago
- [Tech Report] Expanded Hyper-Connections☆66Jul 21, 2026Updated 2 months ago
- [ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.☆123Dec 20, 2024Updated last year
- ☆12Dec 19, 2024Updated last year
- Rad-cGAN v1.0: Radar-based precipitation nowcasting model with conditional Generative Adversarial Networks for multiple dam domains☆11Jul 22, 2022Updated 4 years ago
- ☆19Feb 9, 2026Updated 7 months ago
- The Newton-Muon optimizer☆33Jun 5, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [ICLR 2025] Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models☆163Jul 9, 2025Updated last year
- Simple Python Socket-based Split Learning technique using PyTorch☆14Mar 13, 2020Updated 6 years ago
- ☆38Apr 5, 2024Updated 2 years ago
- [ 공모전 ] 다각적 모델을 활용한 대출 신청 여부 예측과 고객 군집 별 서비스 메시지 제안 : 이상치 탐지, 머신러닝, 딥러닝 모델☆14Jan 28, 2023Updated 3 years ago
- Code for paper "Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion"☆14Mar 28, 2024Updated 2 years ago
- [ICLRW'26] EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation☆50Apr 21, 2026Updated 5 months ago
- Official PyTorch Implementation of EMoE: Unlocking Emergent Modularity in Large Language Models [main conference @ NAACL2024]☆40May 28, 2024Updated 2 years ago
- ☆87Aug 19, 2026Updated last month
- Code for paper "FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning" Neurips2025.☆16Jan 29, 2026Updated 7 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆18Mar 10, 2023Updated 3 years ago
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization☆19Mar 7, 2025Updated last year
- [EMNLP 24] Source code for paper 'AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tu…☆13Dec 15, 2024Updated last year
- Official repository for the ICLR 2026 paper "SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition".☆23Updated this week
- A repository for LotteryFL re-implementation and experiments☆13Dec 18, 2020Updated 5 years ago
- Code for WikiAsp: Multi-document aspect-based summarization.☆43Dec 9, 2020Updated 5 years ago
- NVFP4 Flash-Attention 4 on BlackWell☆65Updated this week
- ☆25Jan 18, 2026Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆10Sep 13, 2022Updated 4 years ago
- ☆21Jul 1, 2024Updated 2 years ago
- ☆23Dec 8, 2022Updated 3 years ago
- Code and data to explore neural scaling laws of xLSTM and Transformer models.☆24Apr 8, 2026Updated 5 months ago
- Expanding linear RNN state-transition matrix eigenvalues to include negatives improves state-tracking tasks and language modeling without…☆22Mar 15, 2025Updated last year
- PELA: Learning Parameter-Efficient Models with Low-Rank Approximation [CVPR 2024]☆19Apr 14, 2024Updated 2 years ago
- ☆20Dec 24, 2024Updated last year