A carefully curated collection of high-quality libraries, projects, tutorials, research papers, and other essential resources focused on Mechanistic Interpretability, a growing subfield in machine learning interpretability research that aims to reverse-engineer neural networks into understandable computational components.
☆127Jul 21, 2026Updated this week
Alternatives and similar repositories for awesome-mechanistic-interpretability
Users that are interested in awesome-mechanistic-interpretability are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A curated reading list for researchers in the Philosophy of Interpretability☆17Aug 17, 2025Updated 11 months ago
- A repository for awesome resources in mechanistic interpretability☆16Jan 18, 2023Updated 3 years ago
- A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository agg…☆215Mar 4, 2026Updated 4 months ago
- Sparse Autoencoders (SAE) vs CLIP fine-tuning fun.☆18Dec 19, 2024Updated last year
- Repository of CFM: Language-aligned Concept Foundation Model for Vision☆21Apr 27, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for the paper "Coherent Concept-based Explanations in Medical Image and Its Application to Skin Lesion Diagnosis", IEEE CVPRW 2023.☆19Dec 13, 2024Updated last year
- The Github repo for our survey paper: "Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large…☆149Apr 15, 2026Updated 3 months ago
- Training Sparse Autoencoders on Language Models☆1,477Updated this week
- Multi-dimensional analysis of orthogonal safety directions in LLM alignment☆22Jun 12, 2026Updated last month
- open source interpretability platform 🧠☆1,075Updated this week
- [ICML 2025 Poster] SAE-V: Interpreting Multimodal Models for Enhanced Alignment☆17Jun 5, 2025Updated last year
- The nnsight package enables interpreting and manipulating the internals of deep learned models.☆995Updated this week
- [CVPR Findings 2026] "Circuit Tracing in Vision-Language Models"☆27Jul 14, 2026Updated last week
- ☆17Feb 9, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Code for Negation Neglect☆16May 22, 2026Updated last month
- Code for "A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models"☆17Jul 20, 2025Updated last year
- Parameter Decomposition☆133Updated this week
- This repository contains the code used for the experiments in the paper "Language Models use Lookbacks to Track Beliefs".☆16Mar 14, 2026Updated 4 months ago
- [NeurIPS2024] Official code for (IMA) Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs☆23Oct 15, 2024Updated last year
- Efficient multi-token attribution for reasoning language models — Python package, CLI, and HTML token traces☆31Updated this week
- ☆10Nov 7, 2022Updated 3 years ago
- ☆95Apr 18, 2026Updated 3 months ago
- The source code of [WWW 2025] MoDiCF☆16Mar 26, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Reasoning Activation in LLMs via Small Model Transfer (NeurIPS 2025)☆22Oct 16, 2025Updated 9 months ago
- A Mechanistic Interpretability Toolkit for Cross-Layer Transcoder Training and Attribution-Graph Visualization☆102Jul 10, 2026Updated last week
- end-to-end dialog system dataset☆13Sep 15, 2019Updated 6 years ago
- ☆12Jul 30, 2025Updated 11 months ago
- Code for the paper: Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery. ECCV 2024.☆59Nov 3, 2024Updated last year
- This is the source code for: Context-aware Entity Typing in Knowledge Graphs.☆16May 10, 2022Updated 4 years ago
- SimpleMagnifyingView can be used as a magnifier as the one iOS providing 🔍. SwiftUI supported!☆13May 4, 2022Updated 4 years ago
- Dependency Parsing as Sequence Labeling with BERT☆13Nov 1, 2020Updated 5 years ago
- Repo of "Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems" (ACL 2026)☆16Apr 27, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ICIAP2022 - Learning Semantics for Visual Place Recognition through Multi-Scale Attention☆16May 10, 2022Updated 4 years ago
- ☆18May 5, 2021Updated 5 years ago
- 🚲 Code and benchmark for our COLM 2025 paper - "Thought Tracing: Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models"☆15Aug 8, 2025Updated 11 months ago
- Training LLMs to Report Their Learned Behaviors☆27Apr 28, 2026Updated 2 months ago
- How do transformer LMs encode relations?☆59Feb 24, 2024Updated 2 years ago
- ☆15Nov 6, 2020Updated 5 years ago
- ☆19Mar 5, 2024Updated 2 years ago