A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.
☆215Mar 4, 2026Updated 4 months ago
Alternatives and similar repositories for Awesome-LMMs-Mechanistic-Interpretability
Users that are interested in Awesome-LMMs-Mechanistic-Interpretability are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- awesome SAE papers☆79May 24, 2025Updated last year
- ☆36Jun 13, 2025Updated last year
- ☆15Jan 20, 2026Updated 6 months ago
- ☆260Nov 22, 2024Updated last year
- awesome papers in LLM interpretability☆624Aug 20, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A curated list of resources for activation engineering☆139Oct 2, 2025Updated 9 months ago
- [ICCV 2025] Auto Interpretation Pipeline and many other functionalities for Multimodal SAE Analysis.☆199Sep 26, 2025Updated 10 months ago
- Training Sparse Autoencoders on Language Models☆1,487Updated this week
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- Sparse autoencoders for vision☆64Jul 22, 2026Updated last week
- [ICML'25] Our study systematically investigates massive values in LLMs' attention mechanisms. First, we observe massive values are concen…☆87Jun 20, 2025Updated last year
- [ICCV25 Oral] Token Activation Map to Visually Explain Multimodal LLMs☆190Dec 14, 2025Updated 7 months ago
- ☆86Nov 5, 2024Updated last year
- [NeurIPS 2025] Official Implementation of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"☆31Jun 4, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A library of visualization tools for the interpretability and hallucination analysis of large vision-language models (LVLMs).☆42May 22, 2025Updated last year
- [preprint] sparsity☆22Updated this week
- Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective (ACL 2024)☆58Oct 28, 2024Updated last year
- 关于LLM和Multimodal LLM的paper list☆60Jun 17, 2026Updated last month
- A carefully curated collection of high-quality libraries, projects, tutorials, research papers, and other essential resources focused on …☆130Updated this week
- The Github repo for our survey paper: "Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large…☆150Apr 15, 2026Updated 3 months ago
- A library for mechanistic interpretability of GPT-style language models☆3,723Updated this week
- A framework that allows you to apply Sparse AutoEncoder on any models☆53Jul 11, 2025Updated last year
- FeatureAlignment = Alignment + Mechanistic Interpretability☆35Mar 8, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2025] Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention☆68Jul 16, 2024Updated 2 years ago
- A repository for awesome resources in mechanistic interpretability☆16Jan 18, 2023Updated 3 years ago
- [ACL 2025] Data and Code for Paper VLSBench: Unveiling Visual Leakage in Multimodal Safety☆62Jul 21, 2025Updated last year
- An awesome repository & A comprehensive survey on interpretability of LLM attention heads.☆412Mar 2, 2025Updated last year
- code for EMNLP 2024 paper: Neuron-Level Knowledge Attribution in Large Language Models☆52Nov 17, 2024Updated last year
- ☆16Nov 14, 2025Updated 8 months ago
- [EMNLP 2025 Demo] Extracting internal representations from vision-language models. Beta version.☆123Apr 25, 2026Updated 3 months ago
- The official repo for "Where do Large Vision-Language Models Look at when Answering Questions?"☆72Jan 7, 2026Updated 6 months ago
- A versatile toolkit for applying Logit Lens to modern large language models (LLMs). Currently supports Llama-3.1-8B and Qwen-2.5-7B, enab…☆174Aug 14, 2025Updated 11 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code for my NeurIPS 2024 ATTRIB paper titled "Attribution Patching Outperforms Automated Circuit Discovery"☆48May 31, 2024Updated 2 years ago
- ViT Prisma is a mechanistic interpretability library for Vision and Video Transformers (ViTs).☆380Jul 23, 2025Updated last year
- [ICLR '25] Official Pytorch implementation of "Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations"☆105Nov 30, 2025Updated 7 months ago
- Mechanistic Interpretability toolkit for Vision-Language-Action models☆20Jul 8, 2026Updated 2 weeks ago
- ☆212Nov 17, 2024Updated last year
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory☆48May 17, 2026Updated 2 months ago
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,497Mar 9, 2026Updated 4 months ago