β941Aug 2, 2026Updated last month
Alternatives and similar repositories for natural_language_autoencoders
Users that are interested in natural_language_autoencoders are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β40May 7, 2026Updated 4 months ago
- open source interpretability platform π§β1,134Updated this week
- β99Apr 18, 2026Updated 4 months ago
- Training Sparse Autoencoders on Language Modelsβ1,529Updated this week
- A toolkit for embedding text datasets with sparse autoencodersβ31Mar 24, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A library for mechanistic interpretability of GPT-style language modelsβ3,870Updated this week
- β2,906Updated this week
- Companion code for the global workspace interpretability paperβ1,893Sep 2, 2026Updated last week
- The nnsight package enables interpreting and manipulating the internals of deep learned models.β1,096Updated this week
- Parameter Decompositionβ142Updated this week
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models β¦β275Updated this week
- β306Apr 13, 2026Updated 4 months ago
- Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsβ462Apr 22, 2026Updated 4 months ago
- Extract residual-stream activations and apply steering vectors (including activation oracles) to any vLLM model during inference.β122Aug 11, 2026Updated last month
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.β19Jun 1, 2026Updated 3 months ago
- β186May 1, 2026Updated 4 months ago
- This was designed for interp researchers who want to do research on or with interp agents to give quality of life improvements and fix β¦β146Feb 8, 2026Updated 7 months ago
- β1,285Updated this week
- ControlArena is a collection of settings, model organisms and protocols - for running control experiments.β233Aug 24, 2026Updated 2 weeks ago
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.β83Sep 1, 2026Updated last week
- β653Jun 19, 2025Updated last year
- Code repo for the model organisms and convergent directions of EM papers.β82Sep 22, 2025Updated 11 months ago
- Code for Negation Neglectβ17May 22, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- An alignment auditing agent capable of quickly exploring alignment hypothesisβ1,332Aug 29, 2026Updated 2 weeks ago
- β33Jul 1, 2026Updated 2 months ago
- β428Aug 21, 2025Updated last year
- β27Jan 7, 2026Updated 8 months ago
- β28Sep 3, 2025Updated last year
- β330Jan 12, 2026Updated 8 months ago
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv Universityβ337Feb 8, 2026Updated 7 months ago
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).β274Feb 27, 2026Updated 6 months ago
- The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away froβ¦β171Jan 20, 2026Updated 7 months ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- β129Updated this week
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"β13Jul 18, 2024Updated 2 years ago
- β114Sep 1, 2026Updated last week
- Mechanistic Interpretability Visualizations using Reactβ367Apr 30, 2026Updated 4 months ago
- β111May 23, 2026Updated 3 months ago
- Steering Llama 2 with Contrastive Activation Additionβ250May 23, 2024Updated 2 years ago
- β33Sep 3, 2026Updated last week