[NeurIPS 2025] Official Implementation of ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding.
☆71Jan 28, 2026Updated 8 months ago
Alternatives and similar repositories for ViSpec
Users that are interested in ViSpec are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆38Jul 21, 2025Updated last year
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆49Apr 16, 2026Updated 5 months ago
- MMSpec: Benchmarking Speculative Decoding for Vision-Language Models☆45Jul 2, 2026Updated 2 months ago
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆30Jul 4, 2026Updated 2 months ago
- Implementation of the paper 'Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance' (EMNLP 2025)☆38Dec 16, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆38Mar 9, 2026Updated 6 months ago
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆54May 12, 2026Updated 4 months ago
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)☆56Mar 14, 2025Updated last year
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,542Feb 20, 2026Updated 7 months ago
- [ICML 2025] CoreMatching: Co-adaptive Sparse Inference Framework for Comprehensive Acceleration of Vision Language Model☆16May 27, 2025Updated last year
- Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding.☆221Mar 18, 2026Updated 6 months ago
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆414Apr 22, 2025Updated last year
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,298Jun 27, 2026Updated 3 months ago
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,194Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2025] PEARL: Parallel Speculative Decoding with Adaptive Draft Length☆172Dec 23, 2025Updated 9 months ago
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 6 months ago
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆179Feb 27, 2026Updated 7 months ago
- [ICCV 2025] Official code for paper: Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs☆85Jul 1, 2025Updated last year
- [ICLR 2026] Official code repository for "⚡️VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration"☆58Jun 17, 2026Updated 3 months ago
- Code for the ICLR 2024 paper Federated Text-driven Prompt Generation for Vision-Language Models (https://openreview.net/forum?id=NW31gAyl…☆27May 7, 2024Updated 2 years ago
- [NeurIPS 2025] HoliTom: Holistic Token Merging for Fast Video Large Language Models☆86Oct 10, 2025Updated 11 months ago
- [EMNLP Main 2026]VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.☆26Jul 20, 2026Updated 2 months ago
- A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Sp…☆208Sep 8, 2026Updated 2 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆69Dec 3, 2024Updated last year
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆56Jul 15, 2025Updated last year
- [NeurIPS 2025] Official code for paper: Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.☆109Sep 20, 2025Updated last year
- Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design☆26May 29, 2025Updated last year
- ☆22Jan 2, 2026Updated 8 months ago
- The official repo of VideoAgentTrek☆59Oct 24, 2025Updated 11 months ago
- Scalable Edge-Assisted Serving Framework for Interactive LLMs [NeurIPS 2025 Spotlight]☆35Nov 14, 2025Updated 10 months ago
- [ECCV 2026🔥] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception…☆68Apr 2, 2026Updated 5 months ago
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Official implementation of the paper: A supervised multi-task framework for joint cryo-ET restoration enabled by generative physical simu…☆19Aug 20, 2026Updated last month
- official code for GliDe with a CaPE☆21Aug 13, 2024Updated 2 years ago
- [EMNLP 2025 main 🔥] Code for "Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More"☆122Oct 12, 2025Updated 11 months ago
- [ICLR 2025] DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference☆56Jun 17, 2025Updated last year
- SPEC-RL: Accelerating On-Policy Reinforcement Learning via Speculative Rollouts☆68Dec 1, 2025Updated 9 months ago
- ☆32Jan 16, 2025Updated last year
- ☆18Apr 7, 2025Updated last year