open source interpretability platform π§
β1,086Jul 17, 2026Updated last week
Alternatives and similar repositories for neuronpedia
Users that are interested in neuronpedia are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Training Sparse Autoencoders on Language Modelsβ1,485Updated this week
- β2,875Jul 18, 2026Updated last week
- A library for mechanistic interpretability of GPT-style language modelsβ3,716Updated this week
- The nnsight package enables interpreting and manipulating the internals of deep learned models.β1,000Updated this week
- Mechanistic Interpretability Visualizations using Reactβ358Apr 30, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β109May 23, 2026Updated 2 months ago
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models β¦β268Updated this week
- β24Feb 13, 2026Updated 5 months ago
- β178May 1, 2026Updated 2 months ago
- β428Aug 21, 2025Updated 11 months ago
- β909Jun 9, 2026Updated last month
- This was designed for interp researchers who want to do research on or with interp agents to give quality of life improvements and fix β¦β146Feb 8, 2026Updated 5 months ago
- β1,190Updated this week
- Parameter Decompositionβ136Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β18Jul 9, 2025Updated last year
- Unified access to Large Language Model modules using NNsightβ116Updated this week
- Sparsify transformers with SAEs and transcodersβ734Updated this week
- β212Nov 17, 2024Updated last year
- Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsβ452Apr 22, 2026Updated 3 months ago
- β49May 27, 2025Updated last year
- Open source interpretability artefacts for R1.β183Apr 21, 2025Updated last year
- β16Nov 14, 2025Updated 8 months ago
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).β267Feb 27, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- An alignment auditing agent capable of quickly exploring alignment hypothesisβ1,274Updated this week
- β223Oct 14, 2025Updated 9 months ago
- Performant framework for training, analyzing and visualizing Sparse Autoencoders (SAEs) and their frontier variants.β223Updated this week
- βοΈ Repository for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.β137Oct 27, 2025Updated 8 months ago
- A toolkit for embedding text datasets with sparse autoencodersβ30Mar 24, 2026Updated 4 months ago
- Companion code for the global workspace interpretability paperβ1,578Updated this week
- ControlArena is a collection of settings, model organisms and protocols - for running control experiments.β213Updated this week
- Stanford NLP Python library for benchmarking the utility of LLM interpretability methodsβ210Mar 12, 2026Updated 4 months ago
- A toolkit for describing model features and intervening on those features to steer behavior.β250Mar 16, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β260Nov 22, 2024Updated last year
- A library for efficient patching and automatic circuit discovery.β99Dec 31, 2025Updated 6 months ago
- Inspect: A framework for large language model evaluationsβ2,411Updated this week
- β597Jul 19, 2024Updated 2 years ago
- Sparsify transformers with cross-layer transcodersβ26Nov 14, 2025Updated 8 months ago
- β95Apr 18, 2026Updated 3 months ago
- The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away froβ¦β158Jan 20, 2026Updated 6 months ago