[ICLR 2025π₯] D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
β27Jul 7, 2025Updated last year
Alternatives and similar repositories for D2O
Users that are interested in D2O are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pytorch implementation of our paper accepted by ICML 2024 -- CaM: Cache Merging for Memory-efficient LLMs Inferenceβ50Jun 19, 2024Updated 2 years ago
- β47Nov 25, 2024Updated last year
- [NAACL 2025π₯] MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inferenceβ22Jun 19, 2025Updated last year
- β39Mar 17, 2025Updated last year
- MMDeepResearch-Bench (MMDR)β32Apr 1, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- β24Jul 7, 2023Updated 3 years ago
- The Official Implementation of Ada-KV [NeurIPS 2025]β139Nov 26, 2025Updated 8 months ago
- Codebase for the ACL 2023 paper: White-Box Multi-Objective Adversarial Attack on Dialogue Generation.β16Dec 8, 2023Updated 2 years ago
- β13Jan 25, 2026Updated 6 months ago
- π° Must-read papers on KV Cache Compression (constantly updating π€).β732Apr 15, 2026Updated 3 months ago
- [EMNLP 2025π₯] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspectiveβ20Jan 7, 2026Updated 7 months ago
- β328Jul 10, 2025Updated last year
- [ICLR2026π₯Oral] SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solvingβ15Feb 26, 2026Updated 5 months ago
- Source code of paper ''KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing''β31Oct 24, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Code repo for "CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs".β17Sep 15, 2024Updated last year
- The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.β54Oct 18, 2024Updated last year
- This is the repo for constructing a comprehensive and rigorous evaluation framework for LLM calibration.β14Apr 9, 2024Updated 2 years ago
- Pytorch implementation for "Compressed Context Memory For Online Language Model Interaction" (ICLR'24)β63Apr 18, 2024Updated 2 years ago
- β30Oct 2, 2025Updated 10 months ago
- Awesome-LLM-KV-Cache: A curated list of πAwesome LLM KV Cache Papers with Codes.β463Jun 17, 2026Updated last month
- Implementation of "PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier"β17Jun 27, 2025Updated last year
- AAAI 2022 paper - Unifying Model Explainability and Robustness for Joint Text Classification and Rationale Extractionβ17Dec 23, 2021Updated 4 years ago
- [ICLR 2025] TidalDecode: A Fast and Accurate LLM Decoding with Position Persistent Sparse Attentionβ57Aug 6, 2025Updated last year
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Game UI Glitch Detection via Bug Understandingβ12Jul 31, 2021Updated 5 years ago
- β15Jan 27, 2026Updated 6 months ago
- [ACL 2026 Findings] Living repository for the survey paper βEfficient Inference for Large Vision-Language Models: Bottlenecks, Techniquesβ¦β26Apr 8, 2026Updated 4 months ago
- β48Oct 16, 2025Updated 9 months ago
- π First survey on Attention Sink in Transformers β 200+ papers on utilization, interpretation, and mitigation.β139Jun 5, 2026Updated 2 months ago
- Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)β12Jun 20, 2025Updated last year
- β88Oct 9, 2024Updated last year
- Code and data for "Impact of Evaluation Methodologies on Code Summarization" in ACL 2022.β10Sep 6, 2022Updated 3 years ago
- Code for the paper "Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers"β18Dec 15, 2020Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Repository for the Q-Filters method (https://arxiv.org/pdf/2503.02812)β34Mar 7, 2025Updated last year
- Uncertainty-Aware Curriculum Learning for Neural Machine Translation (ACL 2020)β11Jun 12, 2020Updated 6 years ago
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.β10May 16, 2024Updated 2 years ago
- LLM KV cache compression made easyβ1,164Updated this week
- β10Apr 29, 2023Updated 3 years ago
- [ICML 2024] CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformersβ34Dec 30, 2024Updated last year
- Code for ICML21 paper "Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation"β12Feb 8, 2023Updated 3 years ago