A list of awesome papers on compression and acceleration of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs).
β17May 12, 2026Updated 3 months ago
Alternatives and similar repositories for Awesome-Efficient-Large-Models
Users that are interested in Awesome-Efficient-Large-Models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2026] The open-source code for FlowCache, including accelerated implementations of the MAGI-1 and Skyreels-V2.β31Apr 24, 2026Updated 4 months ago
- [ECCV 2026π₯] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perceptionβ¦β66Apr 2, 2026Updated 4 months ago
- Benchmarking Audio-Visual Social Interactivity in Omni Modelsβ46May 7, 2026Updated 3 months ago
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantizationβ44Aug 13, 2025Updated last year
- [ICLR 2025] Official implementation of paper "Dynamic Low-Rank Sparse Adaptation for Large Language Models".β25Mar 16, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understandingβ364Aug 5, 2026Updated 3 weeks ago
- ThinK: Thinner Key Cache by Query-Driven Pruningβ30Jun 2, 2026Updated 2 months ago
- Code repo for the paper "SpinQuant LLM quantization with learned rotations"β15Mar 20, 2025Updated last year
- [CVPR 2026] EarlyTom: Early Token Compression Completes Fast Video Understandingβ37Jun 22, 2026Updated 2 months ago
- [CVPR 2026] Official repository for "UniComp: Rethinking Video Compression Through Informational Uniqueness"β27Feb 22, 2026Updated 6 months ago
- β28Dec 7, 2021Updated 4 years ago
- β25Dec 11, 2021Updated 4 years ago
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".β131May 17, 2026Updated 3 months ago
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMsβ89Jan 17, 2026Updated 7 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The Source Code for WebCompassβ25May 2, 2026Updated 3 months ago
- [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Modelsβ105Apr 20, 2026Updated 4 months ago
- Here are some implementations of basic hardware units in RTL language (verilog for now), which can be used for area/power evaluation and β¦β15Aug 25, 2023Updated 3 years ago
- [ICML 2026]β18Jul 4, 2026Updated last month
- Multi Stopwatch for Pythonβ12Sep 28, 2019Updated 6 years ago
- Pytorch implementation of our paper accepted by TPAMI 2023 β Lottery Jackpots Exist in Pre-trained Modelsβ35Jun 19, 2023Updated 3 years ago
- Awesome Video Diffusion Transformers Sparse Attention Papersβ53May 30, 2026Updated 3 months ago
- [ICLR 2026] The official repo of "MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs"β47Jul 3, 2026Updated last month
- Pytorch implementation of our paper accepted by ICML 2023 -- "Bi-directional Masks for Efficient N:M Sparse Training"β14Jun 7, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Lab for Digital Design and Computer Architecture Spring 2022 (252-0028-00L) (ETH).β14Mar 1, 2023Updated 3 years ago
- ETH Computer Architecture - Fall 2020β13Feb 26, 2021Updated 5 years ago
- β26Jul 14, 2026Updated last month
- Cross-Self KV Cache Pruning for Efficient Vision-Language Inferenceβ10Dec 15, 2024Updated last year
- [CVPR 2026 Oral] SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Cachingβ24Jun 5, 2026Updated 2 months ago
- [NeurIPS'24]Efficient and accurate memory saving method towards W4A4 large multi-modal models.β102Jan 3, 2025Updated last year
- Single Cycle and Pipeline CPU of RISC-V Architecture designed for Digital Design and Computer Organization Experiments 2021, NJUβ14Jan 17, 2022Updated 4 years ago
- BESA is a differentiable weight pruning technique for large language models.β17Mar 4, 2024Updated 2 years ago
- Artifact for Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantizationβ18May 9, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- β17May 2, 2024Updated 2 years ago
- Accelerator RTL inspired by VEGETA [HPCA'23] and MicroScopiQ [ISCA'25]β16Nov 11, 2025Updated 9 months ago
- β51May 9, 2026Updated 3 months ago
- [NeurIPS 2022] Make Sharpness-Aware Minimization Stronger: A Sparsified Perturbation Approach -- Official Implementationβ48Jun 29, 2023Updated 3 years ago
- Pytorch code of [CVPR 2023] "NAR-Former: Neural Architecture Representation Learning towards Holistic Attributes Prediction".β12Mar 14, 2023Updated 3 years ago
- β10Jul 27, 2020Updated 6 years ago
- NSCSCC 2023 The Second Prize. TEAM PUA FROM HDU.β14Mar 29, 2025Updated last year