β954Aug 2, 2026Updated 2 months ago
Alternatives and similar repositories for natural_language_autoencoders
Users that are interested in natural_language_autoencoders are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β41May 7, 2026Updated 4 months ago
- open source interpretability platform π§β1,152Updated this week
- β103Sep 26, 2026Updated last week
- Training Sparse Autoencoders on Language Modelsβ1,545Updated this week
- A toolkit for embedding text datasets with sparse autoencodersβ31Mar 24, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A library for mechanistic interpretability of GPT-style language modelsβ3,931Updated this week
- β2,913Sep 20, 2026Updated last week
- Companion code for the global workspace interpretability paperβ2,002Sep 19, 2026Updated 2 weeks ago
- The nnsight package enables interpreting and manipulating the internals of deep learned models.β1,117Updated this week
- Parameter Decompositionβ147Sep 25, 2026Updated last week
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models β¦β276Updated this week
- β315Apr 13, 2026Updated 5 months ago
- Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsβ465Apr 22, 2026Updated 5 months ago
- Extract residual-stream activations and apply steering vectors (including activation oracles) to any vLLM model during inference.β130Updated this week
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.β20Updated this week
- β186May 1, 2026Updated 5 months ago
- This was designed for interp researchers who want to do research on or with interp agents to give quality of life improvements and fix β¦β146Feb 8, 2026Updated 7 months ago
- β1,328Updated this week
- ControlArena is a collection of settings, model organisms and protocols - for running control experiments.β246Updated this week
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.β88Sep 1, 2026Updated last month
- β656Jun 19, 2025Updated last year
- Code repo for the model organisms and convergent directions of EM papers.β86Sep 22, 2025Updated last year
- Code for Negation Neglectβ18May 22, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- An alignment auditing agent capable of quickly exploring alignment hypothesisβ1,356Updated this week
- β33Jul 1, 2026Updated 3 months ago
- β429Aug 21, 2025Updated last year
- β30Jan 7, 2026Updated 8 months ago
- β29Sep 3, 2025Updated last year
- β334Jan 12, 2026Updated 8 months ago
- βοΈ Repository for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.β143Oct 27, 2025Updated 11 months ago
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv Universityβ339Feb 8, 2026Updated 7 months ago
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).β276Feb 27, 2026Updated 7 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away froβ¦β175Jan 20, 2026Updated 8 months ago
- β140Sep 19, 2026Updated 2 weeks ago
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"β13Jul 18, 2024Updated 2 years ago
- β117Updated this week
- Mechanistic Interpretability Visualizations using Reactβ366Apr 30, 2026Updated 5 months ago
- β111May 23, 2026Updated 4 months ago
- Steering Llama 2 with Contrastive Activation Additionβ253May 23, 2024Updated 2 years ago