MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
☆43Jul 2, 2026Updated last month
Alternatives and similar repositories for MMSpec
Users that are interested in MMSpec are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆38Jul 21, 2025Updated last year
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆27Jul 4, 2026Updated last month
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆26Apr 8, 2026Updated 3 months ago
- [NeurIPS 2025] Official Implementation of ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding.☆67Jan 28, 2026Updated 6 months ago
- PhyX: Does Your Model Have the "Wits" for Physical Reasoning?☆54Mar 16, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 2 months ago
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆402Apr 22, 2025Updated last year
- BATCH: Adaptive Batching for Efficient MachineLearning Serving on Serverless Platforms☆11Aug 7, 2021Updated 4 years ago
- Evalution: evolve your LLMs with better evals.☆16Jul 25, 2026Updated last week
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆48Apr 16, 2026Updated 3 months ago
- Source code of paper 'LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval' (WWW 2023)☆22Aug 28, 2023Updated 2 years ago
- ☆32Mar 9, 2026Updated 4 months ago
- Source code of paper 'Open Hierarchical Relation Extraction' (NAACL 2021)☆22Mar 4, 2022Updated 4 years ago
- Must-read papers on Fine-grained Entity Typing☆19Jul 7, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Demo for advanced Java final project in 18-19 1 of Canghong Jin☆25Nov 18, 2018Updated 7 years ago
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,285Jun 27, 2026Updated last month
- ☆15Sep 25, 2025Updated 10 months ago
- Official code repo of SimMLM [ICCV 2025]☆30Dec 1, 2025Updated 8 months ago
- A selective knowledge distillation algorithm for efficient speculative decoders☆39Nov 27, 2025Updated 8 months ago
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆21Mar 1, 2026Updated 5 months ago
- [NAACL 2025] MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning☆23May 31, 2025Updated last year
- [CVPR2026] VecAttention: Vector-wise Sparse Attention for Accelerating Long-Context Inference☆20May 27, 2026Updated 2 months ago
- ☆12Jan 12, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆12Jun 2, 2023Updated 3 years ago
- YOLOv1 implementation of Pytorch with multi-backbone☆10May 8, 2019Updated 7 years ago
- a high performance server framework☆12Dec 11, 2022Updated 3 years ago
- Source code for Jellyfish, a soft real-time inference serving system☆15Dec 20, 2022Updated 3 years ago
- [JRTIP 2023] Efficient Convolutional Neural Networks on Raspberry Pi for Image Classification☆10Aug 12, 2025Updated 11 months ago
- offical implementation of "Calibrating Multimodal Learning" on ICML 2023☆20Jun 5, 2023Updated 3 years ago
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆17Jul 11, 2026Updated 3 weeks ago
- ☆15Apr 11, 2024Updated 2 years ago
- Source code repo for paper "TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation"☆10Aug 11, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [WWW '25] Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability☆18May 30, 2025Updated last year
- Hierarchical Speculative Decoding is the SOTA verification algorithm for lossless accelerated LLM inference.☆25Apr 14, 2026Updated 3 months ago
- Learning MLPs to replace GNN☆10Jun 3, 2023Updated 3 years ago
- Network- and GPU-aware management of serverless functions at the edge☆15Mar 3, 2023Updated 3 years ago
- [MIR] Pytorch Implementation for FM2S, a denoising algorithm for fluorescence microscopy.☆15Mar 13, 2026Updated 4 months ago
- The code for the paper "Efficient Self-Supervised Video Hashing with Selective State Spaces" (AAAI'25).☆24Aug 2, 2025Updated last year
- ☆10Oct 5, 2023Updated 2 years ago