[ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering
☆25Oct 26, 2025Updated 9 months ago
Alternatives and similar repositories for HC-SMoE
Users that are interested in HC-SMoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2026] SERE: Similarity-Based Expert Re-routing for Efficient Batch Decoding in MoE Models☆18Feb 4, 2026Updated 5 months ago
- [ICLR 2022] Denoising Likelihood Score Matching for Conditional Score-based Data Generation☆11Jun 15, 2026Updated last month
- (ICLR 2026) Unveiling Super Experts in Mixture-of-Experts Large Language Models☆44Sep 25, 2025Updated 10 months ago
- D^2-MoE: Delta Decompression for MoE-based LLMs Compression☆82Mar 25, 2025Updated last year
- Official implementation for LaCo (EMNLP 2024 Findings)☆22Oct 3, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- REAM: Merging Improves Pruning of Experts in LLMs☆22Apr 16, 2026Updated 3 months ago
- This repository lists related work using MVC methods for applications.☆18Dec 14, 2023Updated 2 years ago
- [ICML'25] Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models☆23Sep 7, 2025Updated 10 months ago
- MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models☆27May 23, 2026Updated 2 months ago
- MobileNetV2: Inverted Residuals and Linear Bottlenecks☆10Jul 22, 2019Updated 7 years ago
- Efficient Mixture of Experts for LLM Paper List☆184Jun 16, 2026Updated last month
- Code for "Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models?" [ICML 2023]