ThinK: Thinner Key Cache by Query-Driven Pruning
☆30Jun 2, 2026Updated 3 months ago
Alternatives and similar repositories for ThinK
Users that are interested in ThinK are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.☆16May 25, 2025Updated last year
- Make reasoning models scalable☆51Jun 2, 2026Updated 3 months ago
- [ICML 2024] SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models☆22May 28, 2024Updated 2 years ago
- The first spoken long-text dataset derived from live streams, designed to reflect the redundancy-rich and conversational nature of real-w…☆12Jun 28, 2025Updated last year
- ☆15Sep 24, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- VidKV: Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models☆25Mar 26, 2025Updated last year
- [ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.☆57May 2, 2025Updated last year
- [EMNLP 2025🔥] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective☆20Jan 7, 2026Updated 8 months ago
- A list of awesome papers on compression and acceleration of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs).☆18May 12, 2026Updated 4 months ago
- PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation [NeurIPS 2025]☆19Oct 11, 2025Updated 11 months ago
- ☆30Apr 14, 2025Updated last year
- Github Repo for OATS: Outlier-Aware Pruning through Sparse and Low Rank Decomposition☆21Apr 16, 2025Updated last year
- Unofficial implementations of block/layer-wise pruning methods for LLMs.☆78Apr 29, 2024Updated 2 years ago
- [ICLR 2025] The official pytorch implement of "Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Cont…☆72Sep 18, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs☆89Jan 17, 2026Updated 8 months ago
- The Official Implementation of Ada-KV [NeurIPS 2025]☆140Nov 26, 2025Updated 9 months ago
- ☆43Jun 8, 2026Updated 3 months ago
- [ACL Findings 2026] Official Implementation of "FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acc…☆33Apr 14, 2026Updated 5 months ago
- This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"☆38Jul 11, 2024Updated 2 years ago
- InvDesFlow for the inverse design of high-temperature superconducting materials, integrating generative models, stability models, superco…☆21Apr 8, 2026Updated 5 months ago
- ☆12Oct 9, 2023Updated 2 years ago
- ArcLight: A Lightweight LLM Inference Framework☆53May 30, 2026Updated 3 months ago
- An implementation of SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs (NeurIPS 2025)☆34Oct 31, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Mar 28, 2025Updated last year
- ACL 2024: LoRA-Flow Dynamic LoRA Fusion for Large Language Models in Generative Tasks☆25Oct 9, 2024Updated last year
- KV cache compression via sparse coding☆18Oct 26, 2025Updated 10 months ago
- CFG-GAN: Composite functional gradient learning of generative adversarial models☆15Jul 9, 2020Updated 6 years ago
- Extension of libSVM to support Open Set Recognitoin as described in "Toward Open Set Recognition", TPAMI July 2013☆12Oct 21, 2013Updated 12 years ago
- LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning☆39Apr 4, 2024Updated 2 years ago
- Algorithm-System Co-design: accurate and efficient 2-bit KV cache quantization for LLM Inference.☆20May 20, 2026Updated 4 months ago
- [ACL 2024] Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models☆123May 24, 2024Updated 2 years ago
- This is the official repo for "Differentiable Model Scaling using Differentiable Topk"☆12May 16, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs☆221Jul 10, 2026Updated 2 months ago
- AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference☆21Jan 24, 2025Updated last year
- Official Implementation of "Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding" (ICML'25)☆33May 14, 2026Updated 4 months ago
- Code for the paper: CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models☆38Jun 2, 2026Updated 3 months ago
- The official repo of continuous speculative decoding☆36Mar 28, 2025Updated last year
- Pytorch implementation of our paper accepted by ICML 2023 -- "Bi-directional Masks for Efficient N:M Sparse Training"☆14Jun 7, 2023Updated 3 years ago
- This is the model zoo for our CVPR 2023 paper: EfficientViT: Memory Efficient Vision Transformer with Cascaded Group Attention☆14Mar 13, 2024Updated 2 years ago