Official Implementation of "GRIFFIN: Effective Token Alignment for Faster Speculative Decoding"[NeurIPS 2025]
☆19May 12, 2025Updated last year
Alternatives and similar repositories for GRIFFIN
Users that are interested in GRIFFIN are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)☆56Mar 14, 2025Updated last year
- 4-bit Shampoo for Memory-Efficient Network Training (NeurIPS 2024)☆13Feb 13, 2025Updated last year
- (ACL2025 oral) SCOPE: Optimizing KV Cache Compression in Long-context Generation☆36May 28, 2025Updated last year
- Activation-aware Singular Value Decomposition for Compressing Large Language Models☆92Oct 22, 2024Updated last year
- Paper to Reviewer Assignment is a tedious but a very crucial job for conference organizers. Till date the Toronto Paper Matching System (…☆10Nov 30, 2017Updated 8 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.☆10May 16, 2024Updated 2 years ago
- ☆14Jul 23, 2023Updated 3 years ago
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.☆54Oct 18, 2024Updated last year
- ☆18Apr 23, 2025Updated last year
- [NeurIPS 2024] Fast Best-of-N Decoding via Speculative Rejection☆56Oct 29, 2024Updated last year
- ☆47Nov 25, 2024Updated last year
- Official evaluation models and configuration for Stage 1 of the SAIR Mathematics Distillation Challenge: Equational Theories.☆18Apr 19, 2026Updated 3 months ago
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 2 months ago
- A curated collection of resources for building “AI Scientist” systems: AI that assists scientific discovery through literature intelligen…☆15Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.☆56May 2, 2025Updated last year
- TKDE'20 paper, "CONNA: Addressing Name Disambiguation on the Fly".☆15May 31, 2021Updated 5 years ago
- ☆15Jul 30, 2025Updated 11 months ago
- ☆20Jul 23, 2024Updated 2 years ago
- Python trust-region subproblem solvers for nonlinear optimization☆30Updated this week
- ☆19Sep 29, 2024Updated last year
- Trends of arxiv submissions counted from twitter/medium/reddit etc.☆39Dec 11, 2022Updated 3 years ago
- ☆18Aug 17, 2024Updated last year
- [ EMNLP 2025 Main ] Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs☆17Nov 7, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ACL 2025 (Findings)] DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling☆22Dec 16, 2024Updated last year
- Code for our TKDE paper "Understanding WeChat User Preferences and “Wow” Diffusion"☆20Aug 29, 2024Updated last year
- Hierarchical Speculative Decoding is the SOTA verification algorithm for lossless accelerated LLM inference.☆24Apr 14, 2026Updated 3 months ago
- [WSDM 2026] LookAhead Tuning: Safer Language Models via Partial Answer Previews☆17Dec 14, 2025Updated 7 months ago
- A toolkit for modeling and simulation of cloud-native applications.☆16Aug 4, 2025Updated 11 months ago
- Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation, ICML 2024☆24Jun 26, 2024Updated 2 years ago
- Code accompanying our EMNLP 2019 paper: "Revisiting the Evaluation of Theory of Mind through Question Answering"☆29Aug 9, 2020Updated 5 years ago
- [EMNLP'24 (Main)] DRPO(Dynamic Rewarding with Prompt Optimization) is a tuning-free approach for self-alignment. DRPO leverages a search-…☆24Nov 17, 2024Updated last year
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,478Feb 20, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The source code for self-supervised Taxonomy Completion framework TaxoEnrich, published in WWW 2022.☆21Apr 25, 2022Updated 4 years ago
- This is the official repo for Towards Uncertainty-Aware Language Agent.☆31Aug 15, 2024Updated last year
- Code for the EMNLP24 paper "A simple and effective L2 norm based method for KV Cache compression."☆19Dec 13, 2024Updated last year
- [ACL 2026] Repository of IPBench☆23Apr 6, 2026Updated 3 months ago
- Source code for the paper "LongGenBench: Long-context Generation Benchmark"☆24Oct 8, 2024Updated last year
- [AAAI 2024] SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research☆31Aug 6, 2024Updated last year
- A selective knowledge distillation algorithm for efficient speculative decoders☆39Nov 27, 2025Updated 7 months ago