[ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects”.
☆26Apr 8, 2026Updated 4 months ago
Alternatives and similar repositories for Efficient-LVLMs-Inference
Users that are interested in Efficient-LVLMs-Inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆28Jul 4, 2026Updated last month
- ☆19Feb 18, 2025Updated last year
- ☆35Nov 18, 2025Updated 9 months ago
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆48Apr 16, 2026Updated 4 months ago
- ☆15Jan 27, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2026] Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models☆31Mar 21, 2026Updated 5 months ago
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".☆131May 17, 2026Updated 3 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆44Jul 2, 2026Updated last month
- Rookie's guide☆14Aug 10, 2024Updated 2 years ago
- ☆17Mar 24, 2025Updated last year
- The Official Implementation of Ada-KV [NeurIPS 2025]☆139Nov 26, 2025Updated 8 months ago
- ☆28Mar 5, 2026Updated 5 months ago
- ☆33Mar 9, 2026Updated 5 months ago
- ☆36Jun 3, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆11Sep 27, 2022Updated 3 years ago
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs☆89Jan 17, 2026Updated 7 months ago
- [NeurIPS 2024] The official implementation of ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification☆33Mar 30, 2025Updated last year
- ERGO (Efficient Reasoning & Guided Observation) is a large vision-language model trained with reinforcement learning on efficiency object…☆19Feb 25, 2026Updated 5 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 6 months ago
- ☆73Jan 26, 2026Updated 6 months ago
- Myers Research Group's official webpage☆15May 5, 2026Updated 3 months ago
- Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)☆12Jun 20, 2025Updated last year
- [CVPR2026] VecAttention: Vector-wise Sparse Attention for Accelerating Long-Context Inference☆21May 27, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICLR'25] Streaming Video Question-Answering with In-context Video KV-Cache Retrieval☆122Nov 4, 2025Updated 9 months ago
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆112Mar 14, 2025Updated last year
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆224Feb 11, 2026Updated 6 months ago
- Code and data for VTCBench, a VLM benchmark for long-context understanding capabilities under vision-text compression paradigm.☆27Mar 16, 2026Updated 5 months ago
- [AAAI 26'] This is the official pytorch implementation for paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acc…☆46Nov 13, 2025Updated 9 months ago
- nku(Nankai University)南开大学操作系统课程实验 2024Fall☆12Dec 18, 2024Updated last year
- [EMNLP Main 2026]VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.☆26Jul 20, 2026Updated last month
- [Neurips’25] Code for the paper "Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization"☆31Sep 25, 2025Updated 10 months ago
- [ICLR 2025🔥] D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models☆27Jul 7, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 语音合成VITS 纯中文微调☆12Mar 15, 2023Updated 3 years ago
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated 3 weeks ago
- PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding☆30Jun 10, 2026Updated 2 months ago
- Finetuning LLaMA with DeepSpeed☆10Apr 14, 2023Updated 3 years ago
- [NeurIPS 2024] The official implementation of "Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exitin…☆73Jun 26, 2024Updated 2 years ago
- Collection of papers about video-audio understanding☆25Dec 26, 2025Updated 7 months ago
- Implementation of the paper Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs☆12Jun 7, 2025Updated last year