MoE-Visualizer is a tool designed to visualize the selection of experts in Mixture-of-Experts (MoE) models.
☆16Apr 8, 2025Updated last year
Alternatives and similar repositories for MoE-Visualizer
Users that are interested in MoE-Visualizer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pytorch Code for FedHyper☆11Aug 28, 2024Updated 2 years ago
- Code release for AdapMoE accepted by ICCAD 2024☆39Apr 28, 2025Updated last year
- 以帮助你快速找到 LLM 相关工作,尽快抓住 AI 红利为目标的【LLM 教程】☆168Updated this week
- CRAI is a multimodal large language model based on the Mixture of Experts (MoE) architecture, supporting text and image cross-modal tasks…☆16Apr 29, 2025Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆22Oct 3, 2024Updated last year
- Multi-Agent System for Science of Science☆22Feb 5, 2026Updated 6 months ago
- Scaling Laws for Mixture of Experts Models☆15Feb 25, 2025Updated last year
- [ICML 2025] Code for "R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts"☆19Mar 10, 2025Updated last year
- Mixture-of-Experts Multimodal Variational Autoencoder☆15Jul 3, 2025Updated last year
- This System for a Health Plus. The system includes Registration of patients, Making appointments, Storing patient records, Billing in the…☆56Nov 20, 2025Updated 9 months ago
- The official implementation of the paper "Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts" (ICLR 2026).☆22Aug 22, 2026Updated last week
- [WACV 2025] Code for Enhancing Vision-Language Few-Shot Adaptation with Negative Learning☆13Feb 24, 2025Updated last year
- Implementation for ACL 2024 paper "Meta-Task Prompting Elicits Embeddings from Large Language Models"☆12Jul 25, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICLR 2025] Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization☆25Oct 5, 2025Updated 10 months ago
- A curated list of datasets and benchmarks for Vision-Language-Action (VLA) research, with a focus on evaluation protocols and practical g…☆35Apr 28, 2026Updated 4 months ago
- The code for "MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking"☆19Jan 25, 2025Updated last year
- [NAACL 2025] A Closer Look into Mixture-of-Experts in Large Language Models☆60Feb 7, 2025Updated last year
- [ACL 2026 Main] Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis☆50Jun 30, 2026Updated 2 months ago
- Uncertainty-aware Fine-tuning of Segmentation Foundation Models (NeurIPS 2024).☆16Jan 9, 2025Updated last year
- [CVPR 2025] Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Experts☆26Jun 22, 2025Updated last year
- Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning [ICML 2024]☆21May 2, 2024Updated 2 years ago
- A Management System for a Health Care Facility. The system includes Registration of patients, Making appointments, Storing patient record…☆88Nov 20, 2025Updated 9 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [Tool] AutoRec (2015) PyTorch Implementation☆10Mar 1, 2020Updated 6 years ago
- Source code of ACL 2023 Main Conference Paper "PAD-Net: An Efficient Framework for Dynamic Networks".☆14Aug 22, 2026Updated last week
- 🔥 How to efficiently and effectively compress the CoTs or directly generate concise CoTs during inference while maintaining the reasonin…☆65May 22, 2025Updated last year
- Community Implementation of the paper: "Multi-Head Mixture-of-Experts" In PyTorch☆31Updated this week
- Codebase, data and models for hallucination of pruned models☆16Jan 11, 2025Updated last year
- Code for Negative Yields Positive: Unified Dual-Path Adapter for Vision-Language Models☆25Oct 29, 2024Updated last year
- open Source code for propensity evaluation☆19Apr 25, 2026Updated 4 months ago
- [NeurIPS 2025] Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models☆17Aug 19, 2026Updated 2 weeks ago
- Prototyp MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism☆34Apr 4, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Towards Safe LLM with our simple-yet-highly-effective Intention Analysis Prompting☆21Mar 25, 2024Updated 2 years ago
- The official implementation of the paper "Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity".☆17Aug 22, 2026Updated last week
- [ICML 2025 Oral] Mixture of Lookup Experts☆80Dec 3, 2025Updated 9 months ago
- 🎓Automatically Update circult-eda-mlsys-tinyml Papers Daily using Github Actions (Update Every 8th hours)☆10Updated this week
- ☆16Apr 11, 2024Updated 2 years ago
- Official code for "Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping" (ICLR 2025)☆33Oct 25, 2025Updated 10 months ago
- This repository contains the code for the paper "TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back)…☆15Feb 25, 2026Updated 6 months ago