Code to enable layer-level steering in LLMs using sparse auto encoders
☆34Sep 18, 2025Updated 10 months ago
Alternatives and similar repositories for sae-steering
Users that are interested in sae-steering are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2025 Poster] SAE-V: Interpreting Multimodal Models for Enhanced Alignment☆17Jun 5, 2025Updated last year
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- [NAACL'25 Oral] Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering☆83Jun 20, 2026Updated last month
- Experiments with representation engineering☆14Feb 28, 2024Updated 2 years ago
- A TinyStories LM with SAEs and transcoders☆14Apr 3, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆20Jan 19, 2025Updated last year
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"☆10Jun 23, 2024Updated 2 years ago
- [ICML 2025] Unlearning in Diffusion Models using Sparse Autoencoders☆62Oct 16, 2025Updated 9 months ago
- ☆11Dec 4, 2024Updated last year
- Code for my NeurIPS 2024 ATTRIB paper titled "Attribution Patching Outperforms Automated Circuit Discovery"☆48May 31, 2024Updated 2 years ago
- A curated list of resources for activation engineering☆140Oct 2, 2025Updated 9 months ago
- Performant framework for training, analyzing and visualizing Sparse Autoencoders (SAEs) and their frontier variants.☆223Updated this week
- A framework for implementing equivariant DL☆10May 25, 2021Updated 5 years ago
- ☆15Jan 2, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆72Jan 17, 2025Updated last year
- A Java-based framework for combinatorial test input generation, fault characterization and automated test execution.☆12Jan 22, 2024Updated 2 years ago
- Steering Llama 2 with Contrastive Activation Addition☆240May 23, 2024Updated 2 years ago
- ☆595Jul 19, 2024Updated 2 years ago
- A tiny easily hackable implementation of a feature dashboard.☆17Oct 21, 2025Updated 9 months ago
- Code for "Preference Tuning For Toxicity Mitigation Generalizes Across Languages." Paper accepted at Findings of EMNLP 2024☆18Mar 25, 2025Updated last year
- Training Sparse Autoencoders on Language Models☆1,477Updated this week
- ☆16Feb 24, 2022Updated 4 years ago
- Optimizing diffusion for production-ready speeds☆40Jan 10, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- BPE tokenization implemented in Golang 💙☆11Oct 2, 2023Updated 2 years ago
- ☆10Jul 15, 2024Updated 2 years ago
- Code repository for "Eliciting Secret Knowledge from Language Models"☆23Mar 30, 2026Updated 3 months ago
- Extract residual-stream activations and apply steering vectors (including activation oracles) to any vLLM model during inference.☆117Updated this week
- ☆36Jun 13, 2025Updated last year
- A method for steering llms to better follow instructions☆96Jun 10, 2026Updated last month
- FeatureAlignment = Alignment + Mechanistic Interpretability☆35Mar 8, 2025Updated last year
- Röttger et al. (NAACL 2024): "XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models"☆138Feb 24, 2025Updated last year
- ☆18Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The Happy Faces Benchmark☆15Jul 20, 2023Updated 3 years ago
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).☆265Feb 27, 2026Updated 4 months ago
- Self-Teaching Notes on Gradient Leakage Attacks against GPT-2 models.☆14Mar 18, 2024Updated 2 years ago
- Code for our paper "Decomposing The Dark Matter of Sparse Autoencoders"☆23Feb 6, 2025Updated last year
- helper functions for processing and integrating visual language information with Qwen-VL Series Model☆17Aug 30, 2024Updated last year
- [NeurIPS 2024 Spotlight] Code and data for the paper "Finding Transformer Circuits with Edge Pruning".☆70Aug 15, 2025Updated 11 months ago