The code of Advancing Expert Specialization for Better MoE (NeurIPS2025 oral)
☆36Jan 22, 2026Updated 6 months ago
Alternatives and similar repositories for Auxloss-For-Advancing-Expert-Specialization
Users that are interested in Auxloss-For-Advancing-Expert-Specialization are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Apr 16, 2025Updated last year
- Compositional Muon release☆23Jun 5, 2026Updated last month
- Official Pytorch Implementation of "Outlier-weighed Layerwise Sampling for LLM Fine-tuning" by Pengxiang Li, Lu Yin, Xiaowei Gao, Shiwei …☆35Jun 3, 2025Updated last year
- Use the tokenizer in parallel to achieve superior acceleration☆20Mar 21, 2024Updated 2 years ago
- Code for "Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs"☆19Nov 6, 2025Updated 8 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆24Dec 6, 2025Updated 7 months ago
- FA4-based Relative Attention Kernel developed by TML and Colfax☆17Updated this week
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 3 months ago
- Post-Trained MoE Can Skip Half Experts via Self-Distillation☆38May 19, 2026Updated 2 months ago
- [Tech Report] Expanded Hyper-Connections☆47Updated this week
- [ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.☆118Dec 20, 2024Updated last year
- ☆12Dec 19, 2024Updated last year
- Rad-cGAN v1.0: Radar-based precipitation nowcasting model with conditional Generative Adversarial Networks for multiple dam domains☆11Jul 22, 2022Updated 4 years ago
- ☆17Feb 9, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16May 27, 2026Updated last month
- The Newton-Muon optimizer☆30Jun 5, 2026Updated last month
- [ICLR 2025] Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models☆161Jul 9, 2025Updated last year
- Simple Python Socket-based Split Learning technique using PyTorch☆14Mar 13, 2020Updated 6 years ago
- ☆16May 3, 2026Updated 2 months ago
- ☆38Apr 5, 2024Updated 2 years ago
- [ 공모전 ] 다각적 모델을 활용한 대출 신청 여부 예측과 고객 군집 별 서비스 메시지 제안 : 이상치 탐지, 머신러닝, 딥러닝 모델☆14Jan 28, 2023Updated 3 years ago
- Code for paper "Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion"☆14Mar 28, 2024Updated 2 years ago
- [ICLRW'26] EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation☆49Apr 21, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆18Feb 6, 2026Updated 5 months ago
- Official PyTorch Implementation of EMoE: Unlocking Emergent Modularity in Large Language Models [main conference @ NAACL2024]☆39May 28, 2024Updated 2 years ago
- Coco is a proactive co-assistant that connects user workspace with a broader ecosystem of AI agents.☆19Updated this week
- ☆16Nov 5, 2018Updated 7 years ago
- ☆17Jun 9, 2024Updated 2 years ago
- Official repository for Reflecting Reality: Enabling Diffusion Models to Produce Faithful Mirror Reflections☆20Jun 7, 2025Updated last year
- Official repository for the ICLR 2026 paper "SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition".☆17Jul 16, 2026Updated last week
- ☆78May 29, 2026Updated last month
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization☆19Mar 7, 2025Updated last year
- [EMNLP 24] Source code for paper 'AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tu…☆13Dec 15, 2024Updated last year
- "Syntriever: How to Train Your Retriever with Synthetic Data from LLMs" the Nations of the Americas Chapter of the Association for Comput…☆29Mar 5, 2025Updated last year
- A repository for LotteryFL re-implementation and experiments☆13Dec 18, 2020Updated 5 years ago
- ☆34Dec 31, 2025Updated 6 months ago
- NVFP4 Flash-Attention 4 on BlackWell☆30Updated this week
- Code for WikiAsp: Multi-document aspect-based summarization.☆43Dec 9, 2020Updated 5 years ago