A list of awesome papers on compression and acceleration of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs).
β17May 12, 2026Updated 2 months ago
Alternatives and similar repositories for Awesome-Efficient-Large-Models
Users that are interested in Awesome-Efficient-Large-Models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2026] The open-source code for FlowCache, including accelerated implementations of the MAGI-1 and Skyreels-V2.β30Apr 24, 2026Updated 3 months ago
- [ECCV 2026π₯] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perceptionβ¦β65Apr 2, 2026Updated 4 months ago
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantizationβ44Aug 13, 2025Updated 11 months ago
- [ICLR 2025] Official implementation of paper "Dynamic Low-Rank Sparse Adaptation for Large Language Models".β25Mar 16, 2025Updated last year
- [ICML26] Official Repo for WorldCache: Accelerating World Models for Free via Heterogeneous Token Cachingβ40Jul 23, 2026Updated 2 weeks ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understandingβ369Updated this week
- Code repo for the paper "SpinQuant LLM quantization with learned rotations"β15Mar 20, 2025Updated last year
- β28Dec 7, 2021Updated 4 years ago
- β25Dec 11, 2021Updated 4 years ago
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".β130May 17, 2026Updated 2 months ago
- The Source Code for WebCompassβ22May 2, 2026Updated 3 months ago
- [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Modelsβ102Apr 20, 2026Updated 3 months ago
- Here are some implementations of basic hardware units in RTL language (verilog for now), which can be used for area/power evaluation and β¦β15Aug 25, 2023Updated 2 years ago
- [ICML 2026]β18Jul 4, 2026Updated last month
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Multi Stopwatch for Pythonβ12Sep 28, 2019Updated 6 years ago
- Awesome Video Diffusion Transformers Sparse Attention Papersβ46May 30, 2026Updated 2 months ago
- Pytorch implementation of our paper accepted by TPAMI 2023 β Lottery Jackpots Exist in Pre-trained Modelsβ35Jun 19, 2023Updated 3 years ago
- Lab for Digital Design and Computer Architecture Spring 2022 (252-0028-00L) (ETH).β14Mar 1, 2023Updated 3 years ago
- ETH Computer Architecture - Fall 2020β13Feb 26, 2021Updated 5 years ago
- β24Jul 14, 2026Updated 3 weeks ago
- Cross-Self KV Cache Pruning for Efficient Vision-Language Inferenceβ10Dec 15, 2024Updated last year
- [CVPR 2026 Oral] SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Cachingβ24Jun 5, 2026Updated 2 months ago
- [NeurIPS'24]Efficient and accurate memory saving method towards W4A4 large multi-modal models.β102Jan 3, 2025Updated last year
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- BESA is a differentiable weight pruning technique for large language models.β17Mar 4, 2024Updated 2 years ago
- Artifact for Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantizationβ18May 9, 2025Updated last year
- β17May 2, 2024Updated 2 years ago
- Accelerator RTL inspired by VEGETA [HPCA'23] and MicroScopiQ [ISCA'25]β15Nov 11, 2025Updated 8 months ago
- A curated list of recent papers on efficient video attention for video diffusion models, including sparsification, quantization, and cachβ¦β61Oct 27, 2025Updated 9 months ago
- β50May 9, 2026Updated 3 months ago
- Pytorch code of [CVPR 2023] "NAR-Former: Neural Architecture Representation Learning towards Holistic Attributes Prediction".β12Mar 14, 2023Updated 3 years ago
- β10Jul 27, 2020Updated 6 years ago
- A 2-Way Super-Scalar OoO RISC-V Core Based on Intel P6 Microarchitecture.β17Sep 27, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Artifact for "DX100: A Programmable Data Access Accelerator for Indirection (ISCA 2025)" paperβ19Nov 6, 2025Updated 9 months ago
- [EMNLP 2024] Quantize LLM to extremely low-bit, and finetune the quantized LLMsβ16Jul 18, 2024Updated 2 years ago
- XNAS: An effective, modular, and flexible Neural Architecture Search (NAS) framework.β47Jun 29, 2022Updated 4 years ago
- channel pruning for accelerating very deep neural networksβ13Mar 8, 2021Updated 5 years ago
- The code repository of "MBQ: Modality-Balanced Quantization for Large Vision-Language Models"β93Mar 17, 2025Updated last year
- β10Jul 30, 2021Updated 5 years ago
- This is the official Python version of CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Actβ¦β18Oct 25, 2024Updated last year