π First survey on Attention Sink in Transformers β 200+ papers on utilization, interpretation, and mitigation.
β139Jun 5, 2026Updated 2 months ago
Alternatives and similar repositories for Awesome-Attention-Sink
Users that are interested in Awesome-Attention-Sink are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2026π₯Oral] SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solvingβ15Feb 26, 2026Updated 5 months ago
- β15Jan 20, 2026Updated 6 months ago
- The Github repo for our survey paper: "Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Largeβ¦β154Apr 15, 2026Updated 4 months ago
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memoryβ51May 17, 2026Updated 3 months ago
- [EMNLP 2025π₯] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspectiveβ20Jan 7, 2026Updated 7 months ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [ICML'25] Our study systematically investigates massive values in LLMs' attention mechanisms. First, we observe massive values are concenβ¦β87Jun 20, 2025Updated last year
- PhyX: Does Your Model Have the "Wits" for Physical Reasoning?β55Mar 16, 2026Updated 5 months ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalabilityβ16Feb 26, 2026Updated 5 months ago
- [ICLR 2025π₯] D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Modelsβ27Jul 7, 2025Updated last year
- SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolvingβ72Feb 28, 2026Updated 5 months ago
- [CVPR 2026] PromptEnhancer is a prompt-rewriting tool, refining prompts into clearer, structured versions for better image generation.β3,754Updated this week
- Official Repo for SvS: A Self-play with Variational Problem Synthesis strategy for RLVR trainingβ56Dec 13, 2025Updated 8 months ago
- β20Mar 17, 2025Updated last year
- [preprint] sparsityβ23Jul 26, 2026Updated 3 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- β35Jan 20, 2026Updated 6 months ago
- Paper list for the paper "Authorship Attribution in the Era of Large Language Models: Problems, Methodologies, and Challenges (SIGKDD Expβ¦β19May 25, 2026Updated 2 months ago
- AI-powered StartUp Accelerator Engine built with Next.js, LangChain, PostgreSQL + pgvector. Upload, organize, and chat with documents. Inβ¦β885Updated this week
- personal settings for linux tools, including zsh, vim, tmux, pip.β11Dec 2, 2019Updated 6 years ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understandingβ23Jan 25, 2026Updated 6 months ago
- Replace coding puzzles with real-work simulations.β1,907Jul 10, 2026Updated last month
- TVM Documentation in Chinese Simplified / TVM δΈζζζ‘£β3,903May 20, 2026Updated 2 months ago
- β46Jan 4, 2026Updated 7 months ago
- [EMNLP 2024 Findingsπ₯] Official implementation of ": LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inβ¦β103Nov 9, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for the experiments and websites of the paper "Same Task, Different Circuits"β36Jul 21, 2026Updated 3 weeks ago
- Foundations of Medical Large Language Model Learningβ1,992May 27, 2026Updated 2 months ago
- β27Dec 30, 2025Updated 7 months ago
- The code for paper "LLM-Neo: Parameter Efficient Knowledge Distillation for Large Language Models"β15Mar 2, 2025Updated last year
- [NeurIPS 2025π₯]Main source code of SRPO framework.β193Nov 25, 2025Updated 8 months ago
- Klavis AI: MCP integration platforms that let AI agents use tools reliably at any scaleβ5,790Jun 1, 2026Updated 2 months ago
- Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles β unified tools, skillβ¦β5,828Updated this week
- ζΊε·x-agentβ1,088Aug 19, 2025Updated last year
- β25Jan 7, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Align Anything: Training All-modality Model with Feedbackβ4,666Nov 27, 2025Updated 8 months ago
- Source Code for our ICLR'26 paperβ17Feb 22, 2026Updated 5 months ago
- Stanford CoreNLP annotator implementing jMWE for detecting Multi-Word Expressions / collocations