A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.
☆220Mar 4, 2026Updated 6 months ago
Alternatives and similar repositories for Awesome-LMMs-Mechanistic-Interpretability
Users that are interested in Awesome-LMMs-Mechanistic-Interpretability are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- awesome SAE papers☆81May 24, 2025Updated last year
- ☆36Jun 13, 2025Updated last year
- ☆15Jan 20, 2026Updated 8 months ago
- ☆270Nov 22, 2024Updated last year
- awesome papers in LLM interpretability☆628Aug 20, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [ICCV 2025] Auto Interpretation Pipeline and many other functionalities for Multimodal SAE Analysis.☆200Sep 26, 2025Updated last year
- Training Sparse Autoencoders on Language Models☆1,541Updated this week
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- Sparse autoencoders for vision☆67Updated this week
- [ICML'25] Our study systematically investigates massive values in LLMs' attention mechanisms. First, we observe massive values are concen…☆87Jun 20, 2025Updated last year
- [ICCV25 Oral] Token Activation Map to Visually Explain Multimodal LLMs☆189Dec 14, 2025Updated 9 months ago
- ☆87Nov 5, 2024Updated last year
- [NeurIPS 2025] Official Implementation of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"☆31Jun 4, 2026Updated 3 months ago
- [ICLR 2026] Official implementation of the paper "Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs"☆26Mar 3, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A library of visualization tools for the interpretability and hallucination analysis of large vision-language models (LVLMs).☆43May 22, 2025Updated last year
- [NeurIPS'26] In this work, we investigate how Large Language Models (LLMs) adapt their internal representations when encountering inputs …☆25Updated this week
- Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective (ACL 2024)☆59Oct 28, 2024Updated last year
- 关于LLM和Multimodal LLM的paper list☆65Aug 20, 2026Updated last month
- A carefully curated collection of high-quality libraries, projects, tutorials, research papers, and other essential resources focused on …☆156Updated this week
- The Github repo for our survey paper: "Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large…☆157Apr 15, 2026Updated 5 months ago
- A library for mechanistic interpretability of GPT-style language models☆3,911Updated this week
- A framework that allows you to apply Sparse AutoEncoder on any models☆53Jul 11, 2025Updated last year
- FeatureAlignment = Alignment + Mechanistic Interpretability☆35Mar 8, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [CVPR 2025] Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention☆70Jul 16, 2024Updated 2 years ago
- A repository for awesome resources in mechanistic interpretability☆17Jan 18, 2023Updated 3 years ago
- [ACL 2025] Data and Code for Paper VLSBench: Unveiling Visual Leakage in Multimodal Safety☆62Jul 21, 2025Updated last year
- An awesome repository & A comprehensive survey on interpretability of LLM attention heads.☆415Mar 2, 2025Updated last year
- code for EMNLP 2024 paper: Neuron-Level Knowledge Attribution in Large Language Models☆52Nov 17, 2024Updated last year
- ☆16Nov 14, 2025Updated 10 months ago
- A versatile toolkit for applying Logit Lens to modern large language models (LLMs). Currently supports Llama-3.1-8B and Qwen-2.5-7B, enab…☆177Aug 14, 2025Updated last year
- [EMNLP 2025 Demo] Extracting internal representations from vision-language models. Beta version.☆129Apr 25, 2026Updated 5 months ago
- Code for my NeurIPS 2024 ATTRIB paper titled "Attribution Patching Outperforms Automated Circuit Discovery"☆50May 31, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ViT Prisma is a mechanistic interpretability library for Vision and Video Transformers (ViTs).☆390Jul 23, 2025Updated last year
- [ICLR '25] Official Pytorch implementation of "Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations"☆105Nov 30, 2025Updated 9 months ago
- NeurIPS 2026: Mechanistic Interpretability toolkit for Vision-Language-Action models☆25Updated this week
- ☆213Nov 17, 2024Updated last year
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory☆54May 17, 2026Updated 4 months ago
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,509Mar 9, 2026Updated 6 months ago
- Sparse Autoencoders (SAE) vs CLIP fine-tuning fun.☆18Dec 19, 2024Updated last year