Official Implementation of "GRIFFIN: Effective Token Alignment for Faster Speculative Decoding"[NeurIPS 2025]
☆19May 12, 2025Updated last year
Alternatives and similar repositories for GRIFFIN
Users that are interested in GRIFFIN are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)☆56Mar 14, 2025Updated last year
- (ACL2025 oral) SCOPE: Optimizing KV Cache Compression in Long-context Generation☆36May 28, 2025Updated last year
- Activation-aware Singular Value Decomposition for Compressing Large Language Models☆92Oct 22, 2024Updated last year
- Zeroth-Order Fine-Tuning of LLMs in Random Subspaces (ICCV 2025)☆20Nov 22, 2024Updated last year
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.☆10May 16, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.☆55Oct 18, 2024Updated last year
- [NeurIPS 2024] Fast Best-of-N Decoding via Speculative Rejection☆56Oct 29, 2024Updated last year
- 面向大模型的民族文化数据集☆15May 26, 2025Updated last year
- ☆47Nov 25, 2024Updated last year
- Implementation of AdaCQR(COLING 2025)☆15Dec 30, 2024Updated last year
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆54May 12, 2026Updated 4 months ago
- A curated collection of resources for building “AI Scientist” systems: AI that assists scientific discovery through literature intelligen…☆19Aug 4, 2026Updated last month
- [ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.☆57May 2, 2025Updated last year
- [NeurIPS 2025] Official PyTorch implementation for the paper AutoJudge: Judge Decoding Without Manual Annotation☆21Dec 22, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆20Jul 23, 2024Updated 2 years ago
- LongAttn :Selecting Long-context Training Data via Token-level Attention☆15Jul 16, 2025Updated last year
- ☆19Sep 29, 2024Updated last year
- Trends of arxiv submissions counted from twitter/medium/reddit etc.☆39Dec 11, 2022Updated 3 years ago
- ☆18Aug 17, 2024Updated 2 years ago
- ☆19Jun 3, 2024Updated 2 years ago
- ☆32May 24, 2025Updated last year
- Code for our TKDE paper "Understanding WeChat User Preferences and “Wow” Diffusion"☆20Aug 29, 2024Updated 2 years ago
- [WSDM 2026] LookAhead Tuning: Safer Language Models via Partial Answer Previews☆17Dec 14, 2025Updated 9 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A toolkit for modeling and simulation of cloud-native applications.☆16Aug 4, 2025Updated last year
- [ EMNLP 2025 Main ] Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs☆18Nov 7, 2025Updated 10 months ago
- Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation, ICML 2024☆25Jun 26, 2024Updated 2 years ago
- Code accompanying our EMNLP 2019 paper: "Revisiting the Evaluation of Theory of Mind through Question Answering"☆28Aug 9, 2020Updated 6 years ago
- [EMNLP'24 (Main)] DRPO(Dynamic Rewarding with Prompt Optimization) is a tuning-free approach for self-alignment. DRPO leverages a search-…☆24Nov 17, 2024Updated last year
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,540Feb 20, 2026Updated 7 months ago
- The source code for self-supervised Taxonomy Completion framework TaxoEnrich, published in WWW 2022.☆21Apr 25, 2022Updated 4 years ago
- a brief repo about paper research☆15Sep 4, 2024Updated 2 years ago
- This is the official repo for Towards Uncertainty-Aware Language Agent.☆31Aug 15, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for the EMNLP24 paper "A simple and effective L2 norm based method for KV Cache compression."☆19Dec 13, 2024Updated last year
- [ACL 2026] Repository of IPBench☆23Apr 6, 2026Updated 5 months ago
- Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design☆26May 29, 2025Updated last year
- Source code for the paper "LongGenBench: Long-context Generation Benchmark"☆24Oct 8, 2024Updated last year
- A selective knowledge distillation algorithm for efficient speculative decoders☆39Nov 27, 2025Updated 9 months ago
- [NeurIPS 2023] The source code of "Equivariant Spatio-Temporal Attentive Graph Networks to Simulate Physical Dynamics"☆28Aug 6, 2024Updated 2 years ago
- ☆27Aug 23, 2025Updated last year