[ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects”.
☆26Apr 8, 2026Updated 3 months ago
Alternatives and similar repositories for Efficient-LVLMs-Inference
Users that are interested in Efficient-LVLMs-Inference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆27Jul 4, 2026Updated last month
- ☆19Feb 18, 2025Updated last year
- ☆35Nov 18, 2025Updated 8 months ago
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆48Apr 16, 2026Updated 3 months ago
- ☆15Jan 27, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR 2026] Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models☆29Mar 21, 2026Updated 4 months ago
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".☆128May 17, 2026Updated 2 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆43Jul 2, 2026Updated last month
- ☆17Mar 24, 2025Updated last year
- ☆27Mar 5, 2026Updated 4 months ago
- ☆32Mar 9, 2026Updated 4 months ago
- ☆35Jun 3, 2025Updated last year
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs☆88Jan 17, 2026Updated 6 months ago
- Experiments with representation engineering☆14Feb 28, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [NeurIPS 2024] The official implementation of ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification☆33Mar 30, 2025Updated last year
- ERGO (Efficient Reasoning & Guided Observation) is a large vision-language model trained with reinforcement learning on efficiency object…☆19Feb 25, 2026Updated 5 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 6 months ago
- [ACL 25] The Low-cost Long Context Understanding Benchmark for Large Language Models (Outstanding Paper Award)☆23Jul 30, 2025Updated last year
- ☆74Jan 26, 2026Updated 6 months ago
- Myers Research Group's official webpage☆15May 5, 2026Updated 2 months ago
- Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)☆12Jun 20, 2025Updated last year
- [CVPR2026] VecAttention: Vector-wise Sparse Attention for Accelerating Long-Context Inference☆20May 27, 2026Updated 2 months ago
- [ICLR'25] Streaming Video Question-Answering with In-context Video KV-Cache Retrieval☆122Nov 4, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 5 months ago
- Code and data for VTCBench, a VLM benchmark for long-context understanding capabilities under vision-text compression paradigm.☆27Mar 16, 2026Updated 4 months ago
- [AAAI 26'] This is the official pytorch implementation for paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acc…☆47Nov 13, 2025Updated 8 months ago
- nku(Nankai University)南开大学操作系统课程实验 2024Fall☆12Dec 18, 2024Updated last year
- VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.☆26Jul 20, 2026Updated 2 weeks ago
- 语音合成VITS 纯中文微调☆12Mar 15, 2023Updated 3 years ago
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆127Jul 25, 2026Updated last week
- Finetuning LLaMA with DeepSpeed☆10Apr 14, 2023Updated 3 years ago
- [NeurIPS 2024] The official implementation of "Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exitin…☆73Jun 26, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Collection of papers about video-audio understanding☆25Dec 26, 2025Updated 7 months ago
- Implementation of the paper Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs☆12Jun 7, 2025Updated last year
- [TMLR 2026] Survey: https://arxiv.org/pdf/2507.20198☆376Jul 27, 2026Updated last week
- ☆49May 9, 2026Updated 2 months ago
- [ICML 2025] SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity☆76Mar 10, 2026Updated 4 months ago
- V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models☆34Apr 16, 2026Updated 3 months ago
- [ECCV 2026🔥] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception…☆62Apr 2, 2026Updated 4 months ago