β930Aug 2, 2026Updated 3 weeks ago
Alternatives and similar repositories for natural_language_autoencoders
Users that are interested in natural_language_autoencoders are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β39May 7, 2026Updated 3 months ago
- open source interpretability platform π§β1,115Updated this week
- β95Apr 18, 2026Updated 4 months ago
- Training Sparse Autoencoders on Language Modelsβ1,509Aug 10, 2026Updated 2 weeks ago
- A toolkit for embedding text datasets with sparse autoencodersβ30Mar 24, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A library for mechanistic interpretability of GPT-style language modelsβ3,813Updated this week
- β2,896Updated this week
- Companion code for the global workspace interpretability paperβ1,818Aug 4, 2026Updated 2 weeks ago
- The nnsight package enables interpreting and manipulating the internals of deep learned models.β1,045Updated this week
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models β¦β274Updated this week
- Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsβ459Apr 22, 2026Updated 4 months ago
- β300Apr 13, 2026Updated 4 months ago
- Extract residual-stream activations and apply steering vectors (including activation oracles) to any vLLM model during inference.β121Aug 11, 2026Updated last week
- Parameter Decompositionβ139Aug 14, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.β19Jun 1, 2026Updated 2 months ago
- β183May 1, 2026Updated 3 months ago
- This was designed for interp researchers who want to do research on or with interp agents to give quality of life improvements and fix β¦β146Feb 8, 2026Updated 6 months ago
- ControlArena is a collection of settings, model organisms and protocols - for running control experiments.β225Updated this week
- β647Jun 19, 2025Updated last year
- β1,244Updated this week
- Code repo for the model organisms and convergent directions of EM papers.β81Sep 22, 2025Updated 11 months ago
- Code for Negation Neglectβ16May 22, 2026Updated 3 months ago
- β30Jul 1, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- An alignment auditing agent capable of quickly exploring alignment hypothesisβ1,300Updated this week
- β428Aug 21, 2025Updated last year
- β26Jan 7, 2026Updated 7 months ago
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.β80Jul 20, 2026Updated last month
- β28Sep 3, 2025Updated 11 months ago
- The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away froβ¦β166Jan 20, 2026Updated 7 months ago
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv Universityβ337Feb 8, 2026Updated 6 months ago
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).β270Feb 27, 2026Updated 5 months ago
- β122Updated this week
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"β13Jul 18, 2024Updated 2 years ago
- β111Updated this week
- Mechanistic Interpretability Visualizations using Reactβ367Apr 30, 2026Updated 3 months ago
- β110May 23, 2026Updated 3 months ago
- Steering Llama 2 with Contrastive Activation Additionβ249May 23, 2024Updated 2 years ago
- Training LLMs to Report Their Learned Behaviorsβ28Apr 28, 2026Updated 3 months ago
- β32May 29, 2026Updated 2 months ago