A list of awesome papers on compression and acceleration of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs).
β18May 12, 2026Updated 4 months ago
Alternatives and similar repositories for Awesome-Efficient-Large-Models
Users that are interested in Awesome-Efficient-Large-Models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2026] The open-source code for FlowCache, including accelerated implementations of the MAGI-1 and Skyreels-V2.β34Apr 24, 2026Updated 5 months ago
- [ECCV 2026π₯] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perceptionβ¦β68Apr 2, 2026Updated 6 months ago
- Benchmarking Audio-Visual Social Interactivity in Omni Modelsβ46Sep 27, 2026Updated last week
- β¨β¨[AAAI 2026] This is the official implementation of our paper "QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Viβ¦β79Apr 28, 2025Updated last year
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantizationβ45Aug 13, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICLR 2025] Official implementation of paper "Dynamic Low-Rank Sparse Adaptation for Large Language Models".β25Mar 16, 2025Updated last year
- [ICML26] Official Repo for WorldCache: Accelerating World Models for Free via Heterogeneous Token Cachingβ46Jul 23, 2026Updated 2 months ago
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understandingβ363Aug 5, 2026Updated 2 months ago
- ThinK: Thinner Key Cache by Query-Driven Pruningβ30Jun 2, 2026Updated 4 months ago
- Code repo for the paper "SpinQuant LLM quantization with learned rotations"β15Mar 20, 2025Updated last year
- [CVPR 2026] EarlyTom: Early Token Compression Completes Fast Video Understandingβ37Jun 22, 2026Updated 3 months ago
- [CVPR 2026] Official repository for "UniComp: Rethinking Video Compression Through Informational Uniqueness"β27Feb 22, 2026Updated 7 months ago
- β28Dec 7, 2021Updated 4 years ago
- β25Dec 11, 2021Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The Source Code for WebCompassβ27May 2, 2026Updated 5 months ago
- [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Modelsβ109Apr 20, 2026Updated 5 months ago
- Here are some implementations of basic hardware units in RTL language (verilog for now), which can be used for area/power evaluation and β¦β15Aug 25, 2023Updated 3 years ago
- [ICML 2026]β19Jul 4, 2026Updated 3 months ago
- Multi Stopwatch for Pythonβ12Sep 28, 2019Updated 7 years ago
- Pytorch implementation of our paper accepted by TPAMI 2023 β Lottery Jackpots Exist in Pre-trained Modelsβ35Jun 19, 2023Updated 3 years ago
- Awesome Video Diffusion Transformers Sparse Attention Papersβ59May 30, 2026Updated 4 months ago
- [ICLR 2026] The official repo of "MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs"β47Sep 2, 2026Updated last month
- Pytorch implementation of our paper accepted by ICML 2023 -- "Bi-directional Masks for Efficient N:M Sparse Training"β14Jun 7, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Lab for Digital Design and Computer Architecture Spring 2022 (252-0028-00L) (ETH).β14Mar 1, 2023Updated 3 years ago
- ETH Computer Architecture - Fall 2020β13Feb 26, 2021Updated 5 years ago
- [NeurIPS'24]Efficient and accurate memory saving method towards W4A4 large multi-modal models.β105Jan 3, 2025Updated last year
- BESA is a differentiable weight pruning technique for large language models.β17Mar 4, 2024Updated 2 years ago
- Artifact for Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantizationβ20May 9, 2025Updated last year
- β17May 2, 2024Updated 2 years ago
- Accelerator RTL inspired by VEGETA [HPCA'23] and MicroScopiQ [ISCA'25]β16Nov 11, 2025Updated 10 months ago
- β53May 9, 2026Updated 5 months ago
- Pytorch code of [CVPR 2023] "NAR-Former: Neural Architecture Representation Learning towards Holistic Attributes Prediction".β12Mar 14, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- β10Jul 27, 2020Updated 6 years ago
- Artifact for "DX100: A Programmable Data Access Accelerator for Indirection (ISCA 2025)" paperβ20Nov 6, 2025Updated 11 months ago
- NSCSCC 2023 The Second Prize. TEAM PUA FROM HDU.β15Mar 29, 2025Updated last year
- XNAS: An effective, modular, and flexible Neural Architecture Search (NAS) framework.β47Jun 29, 2022Updated 4 years ago
- channel pruning for accelerating very deep neural networksβ13Mar 8, 2021Updated 5 years ago
- The code repository of "MBQ: Modality-Balanced Quantization for Large Vision-Language Models"β97Mar 17, 2025Updated last year
- β10Jul 30, 2021Updated 5 years ago