The Github repo for our survey paper: "Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models"
β157Apr 15, 2026Updated 5 months ago
Alternatives and similar repositories for Awesome-Actionable-MI-Survey
Users that are interested in Awesome-Actionable-MI-Survey are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β15Jan 20, 2026Updated 8 months ago
- π First survey on Attention Sink in Transformers β 200+ papers on utilization, interpretation, and mitigation.β143Jun 5, 2026Updated 3 months ago
- [ACL 2024] Unveiling Linguistic Regions in Large Language Modelsβ34Jun 9, 2024Updated 2 years ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalabilityβ24Sep 20, 2026Updated last week
- [ICLR2026π₯Oral] SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solvingβ15Feb 26, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β21Jun 24, 2024Updated 2 years ago
- SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolvingβ75Feb 28, 2026Updated 6 months ago
- Official implementation of ICLR 2026 paper "LUMINA: Detecting Hallucinations in RAG System with ContextβKnowledge Signals"β19Jan 31, 2026Updated 7 months ago
- A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generationβ16Aug 28, 2025Updated last year
- [ICLR 2025] General-purpose activation steering libraryβ190Sep 18, 2025Updated last year
- awesome papers in LLM interpretabilityβ628Aug 20, 2025Updated last year
- PhyX: Does Your Model Have the "Wits" for Physical Reasoning?β55Mar 16, 2026Updated 6 months ago
- A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggβ¦β220Mar 4, 2026Updated 6 months ago
- Paper List for our ACL 2026 paper "Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architeβ¦β22Apr 23, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A resource repository for representation engineering in large language modelsβ156Nov 14, 2024Updated last year
- β23Dec 17, 2024Updated last year
- β22Mar 3, 2026Updated 6 months ago
- [EMNLP 2025π₯] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspectiveβ20Jan 7, 2026Updated 8 months ago
- code for EMNLP 2024 paper: Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysisβ12Nov 17, 2024Updated last year
- β102Apr 18, 2026Updated 5 months ago
- code for EMNLP 2024 paper: Neuron-Level Knowledge Attribution in Large Language Modelsβ52Nov 17, 2024Updated last year
- A Unified Framework for High-Performance and Extensible LLM Steeringβ300Updated this week
- A carefully curated collection of high-quality libraries, projects, tutorials, research papers, and other essential resources focused on β¦β156Updated this week
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR 2026] The implementation of paper "AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint"β65Nov 20, 2025Updated 10 months ago
- β19Jun 20, 2025Updated last year
- Providing the answer to "How to do patching on all available SAEs on GPT-2?". It is an official repository of the implementation of the pβ¦β14Jan 26, 2025Updated last year
- personal settings for linux tools, including zsh, vim, tmux, pip.β11Dec 2, 2019Updated 6 years ago
- ICLR 2026β28May 13, 2026Updated 4 months ago
- [NAACL'25 Oral] Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineeringβ85Jun 20, 2026Updated 3 months ago
- A library for mechanistic interpretability of GPT-style language modelsβ3,911Updated this week
- β21Sep 25, 2023Updated 3 years ago
- Training Sparse Autoencoders on Language Modelsβ1,541Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"β17Jun 30, 2025Updated last year
- β60Jul 29, 2026Updated last month
- Code for my NeurIPS 2024 ATTRIB paper titled "Attribution Patching Outperforms Automated Circuit Discovery"β50May 31, 2024Updated 2 years ago
- The Official Repository for Paper "HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?"β16May 2, 2026Updated 4 months ago
- Stanford CoreNLP annotator implementing jMWE for detecting Multi-Word Expressions / collocationsβ15Jan 6, 2017Updated 9 years ago
- β116Updated this week
- This repository contains an extension of fairseq for pixel / visual representations of text for machine translation.β37Feb 2, 2024Updated 2 years ago