[ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects”.
☆28Sep 22, 2026Updated last week
Alternatives and similar repositories for Efficient-LVLMs-Inference
Users that are interested in Efficient-LVLMs-Inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆30Jul 4, 2026Updated 3 months ago
- ☆19Feb 18, 2025Updated last year
- ☆35Nov 18, 2025Updated 10 months ago
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆49Apr 16, 2026Updated 5 months ago
- ☆14Jan 27, 2026Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2026] Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models☆31Mar 21, 2026Updated 6 months ago
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".☆138May 17, 2026Updated 4 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆45Jul 2, 2026Updated 3 months ago
- Rookie's guide☆14Aug 10, 2024Updated 2 years ago
- ☆17Mar 24, 2025Updated last year
- ☆38Mar 9, 2026Updated 6 months ago
- ☆36Jun 3, 2025Updated last year
- ☆32Mar 5, 2026Updated 6 months ago
- ☆11Sep 27, 2022Updated 4 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs☆89Jan 17, 2026Updated 8 months ago
- Experiments with representation engineering☆14Feb 28, 2024Updated 2 years ago
- [NeurIPS 2024] The official implementation of ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification☆33Mar 30, 2025Updated last year
- ERGO (Efficient Reasoning & Guided Observation) is a large vision-language model trained with reinforcement learning on efficiency object…☆19Feb 25, 2026Updated 7 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 8 months ago
- ☆72Jan 26, 2026Updated 8 months ago
- [ACL 25] The Low-cost Long Context Understanding Benchmark for Large Language Models (Outstanding Paper Award)☆23Jul 30, 2025Updated last year
- Myers Research Group's official webpage☆15Updated this week
- Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)☆12Jun 20, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [ICLR'25] Streaming Video Question-Answering with In-context Video KV-Cache Retrieval☆129Nov 4, 2025Updated 11 months ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 7 months ago
- ☆14Jul 15, 2025Updated last year
- [EMNLP'26] Code and data for VTCBench, a VLM benchmark for long-context understanding capabilities under vision-text compression paradigm…☆27Aug 27, 2026Updated last month
- [AAAI 26'] This is the official pytorch implementation for paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acc…☆47Nov 13, 2025Updated 10 months ago
- nku(Nankai University)南开大学操作系统课程实验 2024Fall☆12Dec 18, 2024Updated last year
- [EMNLP Main 2026]VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.☆26Jul 20, 2026Updated 2 months ago
- [ICLR 2025🔥] D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models☆27Jul 7, 2025Updated last year
- 语音合成VITS 纯中文微调☆12Mar 15, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated 2 months ago
- Finetuning LLaMA with DeepSpeed☆10Apr 14, 2023Updated 3 years ago
- [NeurIPS 2024] The official implementation of "Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exitin…☆74Jun 26, 2024Updated 2 years ago
- Collection of papers about video-audio understanding☆25Dec 26, 2025Updated 9 months ago
- Code associated with the paper **Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding**☆230Feb 13, 2025Updated last year
- Implementation of the paper Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs☆12Jun 7, 2025Updated last year
- [TMLR 2026] Survey: https://arxiv.org/pdf/2507.20198☆397Jul 27, 2026Updated 2 months ago