[ICLR 2025] Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
☆25Oct 5, 2025Updated 10 months ago
Alternatives and similar repositories for Drop-Upcycling
Users that are interested in Drop-Upcycling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆24Mar 18, 2026Updated 5 months ago
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- [ICCV 2025] EA-ViT: Efficient Adaptation for Elastic Vision Transformer☆27Jul 28, 2025Updated last year
- Implement of 'The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning'☆13Nov 22, 2024Updated last year
- Scaling Laws for Mixture of Experts Models☆15Feb 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code of paper 'Stochastic Layer-Wise Shuffle for Improving Vision Mamba Training'☆21Jun 10, 2025Updated last year
- This repository contains the code for the paper "ST-MoE-BERT: A Spatial-Temporal Mixture-of-Experts Framework for Long-Term Cross-City Mo…☆16Feb 20, 2025Updated last year
- Mamba R1 represents a novel architecture that combines the efficiency of Mamba's state space models with the scalability of Mixture of Ex…☆25Oct 13, 2025Updated 10 months ago
- This repository is the code of paper "Multi-scale Adaptive Task Attention Network for Few-Shot Learning (ICPR-2022)".☆26Jul 4, 2022Updated 4 years ago
- [NAACL 2025] A Closer Look into Mixture-of-Experts in Large Language Models☆60Feb 7, 2025Updated last year
- [ACL 2026 Main] Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis☆50Jun 30, 2026Updated 2 months ago
- This repository is the code of the paper "DiffusionInst: Diffusion Model for Instance Segmentation".☆28Jan 3, 2023Updated 3 years ago
- [CVPR 2025] Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Experts☆26Jun 22, 2025Updated last year
- Code for boomerang distillation enables zero-shot model size interpolation.☆22Jul 10, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pytorch implementation of our paper accepted by ICML 2023 -- "Bi-directional Masks for Efficient N:M Sparse Training"☆14Jun 7, 2023Updated 3 years ago
- Offical implementation of "MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map" (NeurIPS2024 Oral)☆36Jan 18, 2025Updated last year
- Community Implementation of the paper: "Multi-Head Mixture-of-Experts" In PyTorch☆31Aug 3, 2026Updated 3 weeks ago
- 🚀 LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training☆93Dec 3, 2024Updated last year
- Intrinsic Evolution of Embodied Neural Cellular Automata Ecosystems. Parallelized with Taichi and PyTorch☆18Mar 2, 2026Updated 5 months ago
- ☆14Feb 2, 2021Updated 5 years ago
- ☆26Feb 2, 2025Updated last year
- ☆18Aug 19, 2024Updated 2 years ago
- [NeurIPS 2024] Mixture of Experts for Audio-Visual Learning☆25Jan 19, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts☆36Jul 2, 2024Updated 2 years ago
- ☆14Jul 13, 2025Updated last year
- Multi-GPU supported kmeans clustering for cluser-clip☆15Jun 3, 2024Updated 2 years ago
- CLIP-MoE: Mixture of Experts for CLIP☆58Oct 10, 2024Updated last year
- This repository is the code of paper "Multi-level Metric Learning for Few-shot Image Recognition".(ICANN-2022))☆34Dec 13, 2022Updated 3 years ago
- Prototyp MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism☆34Apr 4, 2025Updated last year
- Code for paper "Merging Multi-Task Models via Weight-Ensembling Mixture of Experts"☆33Jun 7, 2024Updated 2 years ago
- [ICLR 2025] Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better☆16Feb 15, 2025Updated last year
- ☆26Mar 26, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICML 2022] "Linearity Grafting: Relaxed Neuron Pruning Helps Certifiable Robustness" by Tianlong Chen*, Huan Zhang*, Zhenyu Zhang, Shiyu…☆16Jun 22, 2022Updated 4 years ago
- Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier (Accepted ECCV 2024)☆10May 6, 2025Updated last year
- Kaggleのshopee コンペのリポジトリ☆11Jun 7, 2021Updated 5 years ago
- Reference implementation of models from Nyonic Model Factory☆12May 13, 2024Updated 2 years ago
- [ICLR 2025, IEEE TPAMI 2026] Mixture Compressor & MC#☆77Feb 12, 2025Updated last year
- Generate synthetic labeled data for extremely low-resource languages using bilingual lexicons.☆20Oct 3, 2024Updated last year
- Official code for "Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping" (ICLR 2025)☆33Oct 25, 2025Updated 10 months ago