[TMLR 2026] LibMoE: A LIBRARY FOR COMPREHENSIVE BENCHMARKING MIXTURE OF EXPERTS IN LARGE LANGUAGE MODELS
β50May 26, 2026Updated last month
Alternatives and similar repositories for LibMoE
Users that are interested in LibMoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] π CodeMMLU Evaluator: A framework for evaluating LM models on CodeMMLU MCQs benchmark.β29Apr 21, 2025Updated last year
- β68Feb 8, 2022Updated 4 years ago
- β22Jul 30, 2023Updated 2 years ago
- Official Release of NeurIPS 2024 paper "Slot State Space Models"β11Mar 22, 2025Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Modelβ13Feb 11, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β26Nov 22, 2025Updated 8 months ago
- [NAACL 2024] Z-GMOT: Zero-shot Generic Multiple Object Trackingβ12May 19, 2026Updated 2 months ago
- [ACL 2026] SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checkingβ76May 26, 2026Updated last month
- [ICML 2025] Code for "R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts"β19Mar 10, 2025Updated last year
- CRAI is a multimodal large language model based on the Mixture of Experts (MoE) architecture, supporting text and image cross-modal tasksβ¦β16Apr 29, 2025Updated last year
- Source codes for paper "BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity".β19Jan 10, 2026Updated 6 months ago
- β15Jan 24, 2025Updated last year
- [NAACL 2025] A Closer Look into Mixture-of-Experts in Large Language Modelsβ61Feb 7, 2025Updated last year
- [CVPR 2023] Bridging the Gap between Model Explanations in Partially Annotated Multi-label Classificationβ21Oct 12, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Dataset and codes for our paper "New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Catβ¦β14Dec 14, 2024Updated last year
- The code of γM4: Multi-Proxy Multi-Gate Mixture of Experts Network for Multiple Instance Learning in Histopathology Image Analysisγβ14Mar 31, 2025Updated last year
- The code for "MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking"β19Jan 25, 2025Updated last year
- Official codes of the 1st place for The NVIDIA AI City Challenge 2023 - Track 2β20Jul 25, 2023Updated 2 years ago
- Scaling Laws for Mixture of Experts Modelsβ15Feb 25, 2025Updated last year
- [ACL 2023 Findings] Emergent Modularity in Pre-trained Transformersβ26Jun 7, 2023Updated 3 years ago
- β14Sep 7, 2022Updated 3 years ago
- Mixture-of-Experts Multimodal Variational Autoencoderβ15Jul 3, 2025Updated last year
- [ACL 2026 Main] Analytical FFN-to-MoE Restructuring via Activation Pattern Analysisβ46Jun 30, 2026Updated 3 weeks ago
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A web app for both Text-based and Visual Question Answering.β13Nov 13, 2023Updated 2 years ago
- Implementation for MomentumSMoEβ19Apr 19, 2025Updated last year
- [NeurIPS 2025] ExGra-Med: Medical Multi-Modal LLM with Extended Context Alignmentβ42Apr 7, 2026Updated 3 months ago
- lanox vim themeβ15Mar 22, 2016Updated 10 years ago
- English-Vietnamese Machine Translation using Transformer (Pytorch)β12Jun 30, 2023Updated 3 years ago
- MoE-Visualizer is a tool designed to visualize the selection of experts in Mixture-of-Experts (MoE) models.β16Apr 8, 2025Updated last year
- My personal blog about AI, ML and DL πβ11Aug 23, 2023Updated 2 years ago
- [ICLR 2025] Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initializationβ25Oct 5, 2025Updated 9 months ago
- Julia package for association rule learningβ13May 18, 2020Updated 6 years ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- MGPATHβ13Oct 15, 2025Updated 9 months ago
- β20Nov 7, 2024Updated last year
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"β17Jun 30, 2025Updated last year
- Beyond Softmax Loss: Intra-Concentration and Inter-Separability Loss for Classification(I2CS)β12Aug 11, 2020Updated 5 years ago
- A highly modular PyTorch framework with a focus on Neural Architecture Search (NAS).β24Dec 3, 2021Updated 4 years ago
- [ACL 2024] Novel reranking method to select the best solutions for code generationβ16Jun 9, 2024Updated 2 years ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year