Official implementation of MAIA, A Multimodal Automated Interpretability Agent
☆113Oct 22, 2025Updated 10 months ago
Alternatives and similar repositories for maia
Users that are interested in maia are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of FIND (NeurIPS '23) Function Interpretation Benchmark and Automated Interpretability Agents☆52Sep 24, 2024Updated last year
- Minimum Description Length probing for neural network representations☆20Jan 28, 2025Updated last year
- Visual Concept Connectome☆15Jun 23, 2024Updated 2 years ago
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆21Jan 19, 2025Updated last year
- Unified access to Large Language Model modules using NNsight☆119Sep 3, 2026Updated last week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆275Updated this week
- Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning☆24Sep 9, 2024Updated 2 years ago
- Code for "Don't trust your eyes: on the (un)reliability of feature visualizations" (ICML 2024)☆33Nov 15, 2023Updated 2 years ago
- ViT Prisma is a mechanistic interpretability library for Vision and Video Transformers (ViTs).☆388Jul 23, 2025Updated last year
- ☆15Jul 12, 2024Updated 2 years ago
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆30Nov 20, 2024Updated last year
- This repo contains the official implementation of ICCV 2025 paper "MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised L…☆22Sep 12, 2025Updated last year
- Attribute statements generated by LLMs to preceding tokens using attention weights.☆29Apr 22, 2025Updated last year
- Official PyTorch Implementation for the "What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-mod…☆20Sep 26, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This repository contains the code used for the experiments in the paper "Language Models use Lookbacks to Track Beliefs".☆17Mar 14, 2026Updated 5 months ago
- A toolkit for describing model features and intervening on those features to steer behavior.☆259Mar 16, 2026Updated 5 months ago
- The AI that helps you achieve your goals☆11Feb 4, 2024Updated 2 years ago
- Code used for "Training Agents to Self-Report Misbehavior"☆18Feb 27, 2026Updated 6 months ago
- A tiny easily hackable implementation of a feature dashboard.☆17Oct 21, 2025Updated 10 months ago
- ☆18May 25, 2023Updated 3 years ago
- 👋 Overcomplete is a Vision-based SAE Toolbox☆151Dec 4, 2025Updated 9 months ago
- James' cookbook of evaluations and finetuning experiments☆35Feb 19, 2026Updated 6 months ago
- [NeurIPS 24] A new training and evaluation framework for learning interpretable deep vision models and benchmarking different interpretab…☆36Jun 5, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ACL 2025 Findings] Official pytorch implementation of "Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vis…☆25Jul 21, 2024Updated 2 years ago
- Repository for "I am a Strange Dataset: Metalinguistic Tests for Language Models"☆46Jan 11, 2024Updated 2 years ago
- Explaining ML models using LLMs☆25Oct 21, 2024Updated last year
- [CVPR 2025] Concept Bottleneck Autoencoder (CB-AE) -- efficiently transform any pretrained (black-box) image generative model into an int…☆21Jul 13, 2026Updated 2 months ago
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- [NeurIPS 2025 MechInterp Workshop - Spotlight] Official implementation of the paper "RelP: Faithful and Efficient Circuit Discovery in La…☆29Nov 3, 2025Updated 10 months ago
- [NeurIPS'25] Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders☆16May 28, 2025Updated last year
- ☆28Apr 1, 2026Updated 5 months ago
- Sparsify transformers with SAEs and transcoders☆739Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆28Sep 3, 2025Updated last year
- [CVPR 2025] Official implementation of the paper "Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottle…☆23Jun 29, 2025Updated last year
- ☆186May 1, 2026Updated 4 months ago
- This is the official implementation of the Concept Discovery Models paper.☆15Aug 27, 2023Updated 3 years ago
- A library for mechanistic interpretability of GPT-style language models☆3,870Updated this week
- Tracking through Containers and Occluders in the Wild (CVPR 2023) - Official Implementation☆41Jun 7, 2024Updated 2 years ago
- Training Sparse Autoencoders on Language Models☆1,529Updated this week