Mixture-of-Basis-Experts for Compressing MoE-based LLMs
☆37Dec 24, 2025Updated 6 months ago
Alternatives and similar repositories for MoBE
Users that are interested in MoBE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- D^2-MoE: Delta Decompression for MoE-based LLMs Compression☆82Mar 25, 2025Updated last year
- ☆24Aug 20, 2025Updated 11 months ago
- [NeurIPS 2025 (spotlight)] HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs☆16Dec 17, 2025Updated 7 months ago
- This is the official code for OThink-R1 project.☆21Jun 19, 2025Updated last year
- [ICLR25] STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs☆20Jun 3, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Gecko Architecture☆16Jan 13, 2026Updated 6 months ago
- ☆56Jul 7, 2025Updated last year
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 9 months ago
- ☆25Feb 12, 2026Updated 5 months ago
- [NeurIPS 2024] Search for Efficient LLMs☆16Jan 16, 2025Updated last year
- ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL (ICLR 2025 Pytorch Code)☆16May 15, 2025Updated last year
- [NeurIPS 2025] Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards☆18Oct 6, 2025Updated 9 months ago
- [ACL 2025] Squeezed Attention: Accelerating Long Prompt LLM Inference☆58Nov 20, 2024Updated last year
- [EMNLP 2025] AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models☆16Apr 29, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆100Nov 22, 2025Updated 7 months ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 8 months ago
- implementation of dualformer☆25Mar 1, 2025Updated last year
- [ICML25] Agentic Compression Benchmark (ACBench)☆17Jul 2, 2025Updated last year
- ☆15Apr 6, 2026Updated 3 months ago
- ☆45Feb 28, 2026Updated 4 months ago
- GitHub Repository for KDD 2022 paper "Saliency-Regularized Deep Multi-Task Learning"☆12Sep 26, 2023Updated 2 years ago
- Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI.☆273Oct 4, 2025Updated 9 months ago
- Intro to using DSPy with Kuzu to enrich the data within the Nobel Laureate mentorship network☆16Sep 16, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆30Aug 5, 2025Updated 11 months ago
- The official implement of "Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings"☆18Dec 5, 2024Updated last year
- LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering☆22Jun 2, 2026Updated last month
- [ICML 2025] SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models☆62Aug 9, 2024Updated last year
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated 11 months ago
- ☆13Dec 17, 2021Updated 4 years ago
- (ICLR 2026) Unveiling Super Experts in Mixture-of-Experts Large Language Models☆44Sep 25, 2025Updated 9 months ago
- Spectral Tensor Train Parameterization of Deep Learning Layers☆17Jul 1, 2021Updated 5 years ago
- ☆22Nov 26, 2025Updated 7 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated last year
- Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI, derived from Ling.☆109Aug 5, 2025Updated 11 months ago
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. 🚀 The official implementation of https://arx…☆31Feb 17, 2025Updated last year
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated 10 months ago
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆19Mar 1, 2026Updated 4 months ago
- The raw UserRL repo under construction☆111Jun 2, 2026Updated last month
- Awesome LLM pruning papers all-in-one repository with integrating all useful resources and insights.☆172May 19, 2026Updated 2 months ago