[EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
☆48Apr 16, 2026Updated 4 months ago
Alternatives and similar repositories for SpecVLM
Users that are interested in SpecVLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".☆130May 17, 2026Updated 3 months ago
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆26Apr 8, 2026Updated 4 months ago
- [NeurIPS 2025] Official Implementation of ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding.☆70Jan 28, 2026Updated 6 months ago
- ☆19Feb 18, 2025Updated last year
- [AAAI 2024] MLNet: Mutual Learning Network with Neighborhood Invariance for Universal Domain Adaptation☆21Feb 29, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Implementation of our paper "RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO"☆54Jul 11, 2026Updated last month
- [ICLR 2025] Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models☆75Mar 29, 2025Updated last year
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 3 months ago
- ☆15Jan 27, 2026Updated 6 months ago
- Training Code for ADMIS Teams in CVPR2024 FRCSyn Competition☆33Jan 5, 2026Updated 7 months ago
- Official implementation of "EponaV2: Driving World Model with Comprehensive Future Reasoning"☆34May 14, 2026Updated 3 months ago
- The official code of "Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling"☆67Jun 30, 2026Updated last month
- [NeurIPS 2025] HoliTom: Holistic Token Merging for Fast Video Large Language Models☆84Oct 10, 2025Updated 10 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆45Jul 2, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- The Official Implementation of Ada-KV [NeurIPS 2025]☆139Nov 26, 2025Updated 8 months ago
- Rookie's guide☆14Aug 10, 2024Updated 2 years ago
- ☆37Feb 12, 2026Updated 6 months ago
- The official implementation of MotionGrasp☆38Nov 15, 2025Updated 9 months ago
- Official code for **Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity** (PruneSI…☆13Mar 25, 2026Updated 4 months ago
- [ICCV 2025] Official code for paper: Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs☆84Jul 1, 2025Updated last year
- [NeurIPS 2025] FastVID: Dynamic Density Pruning for Fast Video Large Language Models☆37Nov 10, 2025Updated 9 months ago
- [ICML 2025] CoreMatching: Co-adaptive Sparse Inference Framework for Comprehensive Acceleration of Vision Language Model☆16May 27, 2025Updated last year
- [ACL 2026] Enabling Efficient Reasoning in LLMs via Black-box Persuasive Prompting☆22Jan 9, 2026Updated 7 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Code associated with the paper **Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding**☆230Feb 13, 2025Updated last year
- [ACM MM 2025] TimeChat-online: 80% Visual Tokens are Naturally Redundant in Streaming Videos☆133Jun 29, 2026Updated last month
- ☆17Mar 24, 2025Updated last year
- ☆36Jun 3, 2025Updated last year
- Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆132Jul 25, 2026Updated 3 weeks ago
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 4 months ago
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs☆88Jan 17, 2026Updated 7 months ago
- [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs☆61Feb 2, 2026Updated 6 months ago
- [ICLR 2026] Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision☆236May 31, 2026Updated 2 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆19Mar 29, 2026Updated 4 months ago
- MR. Video: MapReduce is the Principle for Long Video Understanding☆31Jun 18, 2026Updated 2 months ago
- official code for GliDe with a CaPE☆21Aug 13, 2024Updated 2 years ago
- ☆11Sep 27, 2022Updated 3 years ago
- [NAACL 2025🔥] MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference☆22Jun 19, 2025Updated last year
- [ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".☆272Jul 30, 2026Updated 2 weeks ago
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆86Jul 14, 2025Updated last year