Training Sparse Autoencoders on Language Models
☆1,501Aug 10, 2026Updated this week
Alternatives and similar repositories for SAELens
Users that are interested in SAELens are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A library for mechanistic interpretability of GPT-style language models☆3,794Updated this week
- Sparsify transformers with SAEs and transcoders☆734Updated this week
- ☆429Aug 21, 2025Updated 11 months ago
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).☆269Feb 27, 2026Updated 5 months ago
- The nnsight package enables interpreting and manipulating the internals of deep learned models.☆1,025Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆182May 1, 2026Updated 3 months ago
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆272Updated this week
- Mechanistic Interpretability Visualizations using React☆363Apr 30, 2026Updated 3 months ago
- ☆110May 23, 2026Updated 2 months ago
- ☆212Nov 17, 2024Updated last year
- Sparse Autoencoder for Mechanistic Interpretability☆303Jul 20, 2024Updated 2 years ago
- Performant framework for training, analyzing and visualizing Sparse Autoencoders (SAEs) and their frontier variants.☆226Updated this week
- ☆599Jul 19, 2024Updated 2 years ago
- ☆224Oct 14, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Using sparse coding to find distributed representations used by neural networks.☆309Nov 10, 2023Updated 2 years ago
- open source interpretability platform 🧠☆1,103Updated this week
- ☆2,895Jul 18, 2026Updated 3 weeks ago
- ☆1,227Updated this week
- ViT Prisma is a mechanistic interpretability library for Vision and Video Transformers (ViTs).☆383Jul 23, 2025Updated last year
- ☆141Oct 28, 2023Updated 2 years ago
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- ☆72Jan 17, 2025Updated last year
- ☆294Oct 1, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆61Nov 19, 2024Updated last year
- Stanford NLP Python library for understanding and improving PyTorch models via interventions☆898Mar 6, 2026Updated 5 months ago
- Implementation of the BatchTopK activation function for training sparse autoencoders (SAEs)☆67Jul 24, 2025Updated last year
- ☆31Apr 4, 2024Updated 2 years ago
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆29Nov 20, 2024Updated last year
- Unified access to Large Language Model modules using NNsight☆119Jul 28, 2026Updated 2 weeks ago
- Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".☆428Jun 13, 2025Updated last year
- Steering Llama 2 with Contrastive Activation Addition☆249May 23, 2024Updated 2 years ago
- Tools for understanding how transformer predictions are built layer-by-layer☆608Aug 7, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆98Mar 28, 2025Updated last year
- A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc..☆308Jan 22, 2026Updated 6 months ago
- ☆78Mar 6, 2025Updated last year
- ☆263Nov 22, 2024Updated last year
- This repository collects all relevant resources about interpretability in LLMs☆404Nov 1, 2024Updated last year
- Parameter Decomposition☆137Updated this week
- ☆26Aug 23, 2025Updated 11 months ago