MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
☆44Jul 2, 2026Updated last month
Alternatives and similar repositories for MMSpec
Users that are interested in MMSpec are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆38Jul 21, 2025Updated last year
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆29Jul 4, 2026Updated last month
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆26Apr 8, 2026Updated 4 months ago
- [NeurIPS 2025] Official Implementation of ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding.☆70Jan 28, 2026Updated 6 months ago
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆405Apr 22, 2025Updated last year
- BATCH: Adaptive Batching for Efficient MachineLearning Serving on Serverless Platforms☆11Aug 7, 2021Updated 5 years ago
- Evalution: evolve your LLMs with better evals.☆16Updated this week
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆48Apr 16, 2026Updated 4 months ago
- Source code of paper 'LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval' (WWW 2023)☆22Aug 28, 2023Updated 2 years ago
- ☆34Mar 9, 2026Updated 5 months ago
- Source code of paper 'Open Hierarchical Relation Extraction' (NAACL 2021)☆22Mar 4, 2022Updated 4 years ago
- [NAACL 2025🔥] MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference☆22Jun 19, 2025Updated last year
- Must-read papers on Fine-grained Entity Typing☆19Jul 7, 2022Updated 4 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,292Jun 27, 2026Updated last month
- ☆15Sep 25, 2025Updated 11 months ago
- Official code repo of SimMLM [ICCV 2025]☆30Dec 1, 2025Updated 8 months ago
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆22Mar 1, 2026Updated 5 months ago
- [ACL'23 Findings] "Aligning Instruction Tasks Unlocks Large Language Models as Zero-Shot Relation Extractors"☆38Dec 22, 2023Updated 2 years ago
- ☆12Jan 12, 2024Updated 2 years ago
- YOLOv1 implementation of Pytorch with multi-backbone☆10May 8, 2019Updated 7 years ago
- a high performance server framework☆12Dec 11, 2022Updated 3 years ago
- ☆20Jan 27, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Collection of memory microbenchmarks to investigate NVIDIA GPUs Network on Chip architectures☆15Apr 14, 2026Updated 4 months ago
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆21Jul 11, 2026Updated last month
- ☆19Feb 18, 2025Updated last year
- ☆15Apr 11, 2024Updated 2 years ago
- [WWW '25] Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability☆18May 30, 2025Updated last year
- ☆12Mar 20, 2023Updated 3 years ago
- Hierarchical Speculative Decoding is the SOTA verification algorithm for lossless accelerated LLM inference.☆25Apr 14, 2026Updated 4 months ago
- Learning MLPs to replace GNN☆10Jun 3, 2023Updated 3 years ago
- Network- and GPU-aware management of serverless functions at the edge☆15Mar 3, 2023Updated 3 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Implementation for paper "Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models".☆131May 17, 2026Updated 3 months ago
- The code for the paper "Efficient Self-Supervised Video Hashing with Selective State Spaces" (AAAI'25).☆24Aug 2, 2025Updated last year
- ☆10Oct 5, 2023Updated 2 years ago
- PyTorch code for full quantization of DNN using BCGD☆14Jul 24, 2019Updated 7 years ago
- [NLPCC 2021] Shared Task on AutoIE2: Sub-Event Identification☆14Jul 19, 2021Updated 5 years ago
- [EMNLP 2023] Question Answering as Programming for Solving Time-Sensitive Questions☆12Dec 18, 2023Updated 2 years ago
- Object as a Service (OaaS)☆16Dec 4, 2025Updated 8 months ago