MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
☆46Jul 2, 2026Updated 2 months ago
Alternatives and similar repositories for MMSpec
Users that are interested in MMSpec are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆38Jul 21, 2025Updated last year
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆29Jul 4, 2026Updated 2 months ago
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆28Apr 8, 2026Updated 5 months ago
- [NeurIPS 2025] Official Implementation of ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding.☆71Jan 28, 2026Updated 7 months ago
- PhyX: Does Your Model Have the "Wits" for Physical Reasoning?☆55Mar 16, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆54May 12, 2026Updated 4 months ago
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆412Apr 22, 2025Updated last year
- BATCH: Adaptive Batching for Efficient MachineLearning Serving on Serverless Platforms☆11Aug 7, 2021Updated 5 years ago
- Evalution: evolve your LLMs with better evals.☆16Sep 8, 2026Updated last week
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆48Apr 16, 2026Updated 4 months ago
- ☆36Mar 9, 2026Updated 6 months ago
- [NAACL 2025🔥] MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference☆22Jun 19, 2025Updated last year
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,296Jun 27, 2026Updated 2 months ago
- ☆17Sep 25, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A selective knowledge distillation algorithm for efficient speculative decoders☆39Nov 27, 2025Updated 9 months ago
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆23Mar 1, 2026Updated 6 months ago
- [NAACL 2025] MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning☆23May 31, 2025Updated last year
- [CVPR2026] VecAttention: Vector-wise Sparse Attention for Accelerating Long-Context Inference☆21May 27, 2026Updated 3 months ago
- ☆12Jan 12, 2024Updated 2 years ago
- YOLOv1 implementation of Pytorch with multi-backbone☆10May 8, 2019Updated 7 years ago
- a high performance server framework☆12Dec 11, 2022Updated 3 years ago
- offical implementation of "Calibrating Multimodal Learning" on ICML 2023☆20Jun 5, 2023Updated 3 years ago
- Source code repo for paper "TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation"☆10Aug 11, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [WWW '25] Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability☆18May 30, 2025Updated last year
- Hierarchical Speculative Decoding is the SOTA verification algorithm for lossless accelerated LLM inference.☆25Apr 14, 2026Updated 5 months ago
- Learning MLPs to replace GNN☆10Jun 3, 2023Updated 3 years ago
- Network- and GPU-aware management of serverless functions at the edge☆15Mar 3, 2023Updated 3 years ago
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".☆133May 17, 2026Updated 3 months ago
- The code for the paper "Efficient Self-Supervised Video Hashing with Selective State Spaces" (AAAI'25).☆24Aug 2, 2025Updated last year
- MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following☆16Oct 31, 2024Updated last year
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).