Mixture-of-Basis-Experts for Compressing MoE-based LLMs
☆37Dec 24, 2025Updated 7 months ago
Alternatives and similar repositories for MoBE
Users that are interested in MoBE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of the ICLR paper "Streamlining Redundant Layers to Compress Large Language Models"☆44May 1, 2025Updated last year
- D^2-MoE: Delta Decompression for MoE-based LLMs Compression☆83Mar 25, 2025Updated last year
- ☆25Aug 20, 2025Updated 11 months ago
- [NeurIPS 2025 (spotlight)] HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs☆16Dec 17, 2025Updated 7 months ago
- This is the official code for OThink-R1 project.☆21Jun 19, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICLR25] STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs☆20Jun 3, 2025Updated last year
- An implementation is provided here for the NeurIPS2024 paper "MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected…☆16Mar 24, 2026Updated 4 months ago
- Gecko Architecture☆18Jan 13, 2026Updated 6 months ago
- ☆56Jul 7, 2025Updated last year
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 9 months ago
- ☆25Feb 12, 2026Updated 5 months ago
- [NeurIPS 2024] Search for Efficient LLMs☆16Jan 16, 2025Updated last year
- ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL (ICLR 2025 Pytorch Code)☆16May 15, 2025Updated last year
- [NeurIPS 2025] Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards☆18Oct 6, 2025Updated 10 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACL 2025] Squeezed Attention: Accelerating Long Prompt LLM Inference☆58Nov 20, 2024Updated last year
- [EMNLP 2025] AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models☆16Apr 29, 2026Updated 3 months ago
- ☆102Nov 22, 2025Updated 8 months ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 9 months ago
- implementation of dualformer☆25Mar 1, 2025Updated last year
- This repository provides the official implementation of QSVD, a method for efficient low-rank approximation that unifies Query-Key-Value …☆28May 16, 2026Updated 2 months ago
- [ICML25] Agentic Compression Benchmark (ACBench)☆18Jul 2, 2025Updated last year
- ☆15Apr 6, 2026Updated 4 months ago
- A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Reports☆17Nov 14, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated last year
- LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering☆22Jun 2, 2026Updated 2 months ago
- Intro to using DSPy with Kuzu to enrich the data within the Nobel Laureate mentorship network☆16Sep 16, 2025Updated 10 months ago
- [ICML 2025] SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models☆63Aug 9, 2024Updated 2 years ago
- ☆13Dec 17, 2021Updated 4 years ago
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated 11 months ago
- (ICLR 2026) Unveiling Super Experts in Mixture-of-Experts Large Language Models☆44Sep 25, 2025Updated 10 months ago
- ☆23Nov 26, 2025Updated 8 months ago
- ABench is an evolving open-source benchmark suite designed to rigorously evaluate and enhance Large Language Models (LLMs) on complex cro…☆28Jul 30, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI, derived from Ling.☆109Aug 5, 2025Updated last year
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. 🚀 The official implementation of https://arx…☆31Feb 17, 2025Updated last year
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated 11 months ago
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆22Mar 1, 2026Updated 5 months ago
- Awesome LLM pruning papers all-in-one repository with integrating all useful resources and insights.☆175May 19, 2026Updated 2 months ago
- 🚀 LLM-I: Transform LLMs into natural interleaved multimodal creators! ✨ Tool-use framework supporting image search, generation, code ex…☆41Oct 20, 2025Updated 9 months ago
- ☆16Apr 6, 2023Updated 3 years ago