☆141Oct 28, 2023Updated 2 years ago
Alternatives and similar repositories for 1L-Sparse-Autoencoder
Users that are interested in 1L-Sparse-Autoencoder are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Using sparse coding to find distributed representations used by neural networks.☆313Nov 10, 2023Updated 2 years ago
- Sparse Autoencoder for Mechanistic Interpretability☆306Jul 20, 2024Updated 2 years ago
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).☆275Feb 27, 2026Updated 6 months ago
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"☆13Jul 18, 2024Updated 2 years ago
- ☆428Aug 21, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- An Open Source Implementation of Anthropic's Paper: "Towards Monosemanticity: Decomposing Language Models with Dictionary Learning"☆69May 12, 2024Updated 2 years ago
- ☆62Nov 19, 2024Updated last year
- A very hacky set of functions for getting plotly to do what I want when doing mech interp research, designed to be compatible with PyTorc…☆15Jun 16, 2023Updated 3 years ago
- ☆225Oct 14, 2025Updated 11 months ago
- Training Sparse Autoencoders on Language Models☆1,538Updated this week
- ☆26Dec 20, 2023Updated 2 years ago
- ☆602Jul 19, 2024Updated 2 years ago
- Mechanistic Interpretability Visualizations using React☆366Apr 30, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Sparsify transformers with SAEs and transcoders☆739Sep 14, 2026Updated last week
- ☆15Jul 12, 2024Updated 2 years ago
- Experiments with representation engineering☆14Feb 28, 2024Updated 2 years ago
- A library for mechanistic interpretability of GPT-style language models☆3,893Updated this week
- ☆297Oct 1, 2024Updated last year
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆275Sep 14, 2026Updated last week
- ☆30Aug 2, 2024Updated 2 years ago
- ☆116Aug 8, 2024Updated 2 years ago
- Code for reproducing our paper "Not All Language Model Features Are Linear"☆94Nov 27, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆24Updated this week
- Tools for exploring Transformer neuron behaviour, including input pruning and diversification.☆11Jun 6, 2023Updated 3 years ago
- ☆25Sep 5, 2024Updated 2 years ago
- Stanford NLP Python library for understanding and improving PyTorch models via interventions☆899Mar 6, 2026Updated 6 months ago
- The nnsight package enables interpreting and manipulating the internals of deep learned models.☆1,105Updated this week
- Repository for the "Chain-of-Thought Reasoning In The Wild Is Not Always Faithful" paper☆35Mar 31, 2026Updated 5 months ago
- Notebooks accompanying Anthropic's "Toy Models of Superposition" paper☆158Sep 14, 2022Updated 4 years ago
- ☆158Updated this week
- ☆1,080Mar 6, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- CausalGym: Benchmarking causal interpretability methods on linguistic tasks☆54Nov 30, 2024Updated last year
- Modified to support crosscoder training.☆29Jul 2, 2026Updated 2 months ago
- ☆1,309Updated this week
- ☆17Feb 14, 2024Updated 2 years ago
- Implementation of path patching & activation patching (will eventually add to TransformerLens).☆15Jan 8, 2024Updated 2 years ago
- Code for Negation Neglect☆18May 22, 2026Updated 3 months ago
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆30Nov 20, 2024Updated last year