A list of awesome papers on compression and acceleration of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs).
β16May 12, 2026Updated 2 months ago
Alternatives and similar repositories for Awesome-Efficient-Large-Models
Users that are interested in Awesome-Efficient-Large-Models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2026] The open-source code for FlowCache, including accelerated implementations of the MAGI-1 and Skyreels-V2.β29Apr 24, 2026Updated 2 months ago
- [ECCV 2026π₯] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perceptionβ¦β63Apr 2, 2026Updated 3 months ago
- Benchmarking Audio-Visual Social Interactivity in Omni Modelsβ46May 7, 2026Updated 2 months ago
- β¨β¨[AAAI 2026] This is the official implementation of our paper "QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Viβ¦β79Apr 28, 2025Updated last year
- [ICLR 2025] Official implementation of paper "Dynamic Low-Rank Sparse Adaptation for Large Language Models".β25Mar 16, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantizationβ44Aug 13, 2025Updated 11 months ago
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understandingβ369May 24, 2026Updated last month
- Code repo for the paper "SpinQuant LLM quantization with learned rotations"β15Mar 20, 2025Updated last year
- [CVPR 2026] Official repository for "UniComp: Rethinking Video Compression Through Informational Uniqueness"β27Feb 22, 2026Updated 4 months ago
- [CVPR 2026] EarlyTom: Early Token Compression Completes Fast Video Understandingβ31Jun 22, 2026Updated 3 weeks ago
- β28Dec 7, 2021Updated 4 years ago
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMsβ88Jan 17, 2026Updated 6 months ago
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".β118May 17, 2026Updated 2 months ago
- β25Dec 11, 2021Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Modelsβ99Apr 20, 2026Updated 3 months ago
- Here are some implementations of basic hardware units in RTL language (verilog for now), which can be used for area/power evaluation and β¦β15Aug 25, 2023Updated 2 years ago
- [ICML 2026]