Training LLMs to Report Their Learned Behaviors
☆29Apr 28, 2026Updated 4 months ago
Alternatives and similar repositories for introspection-adapters
Users that are interested in introspection-adapters are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for Learning to Interpret Weight Differences in Language Models (Goel et al. 2025)☆22Jan 4, 2026Updated 8 months ago
- Codebase for Temporal SAEs paper☆26Nov 14, 2025Updated 9 months ago
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆22Dec 14, 2024Updated last year
- Code and data for the paper "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs"☆50Mar 31, 2026Updated 5 months ago
- Code for Negation Neglect☆17May 22, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [NeurIPS'25] Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders☆16May 28, 2025Updated last year
- Repository for "Training Language Models To Explain Their Own Computations"☆37Jul 7, 2026Updated 2 months ago
- James' cookbook of evaluations and finetuning experiments☆35Feb 19, 2026Updated 6 months ago
- ☆16Oct 13, 2025Updated 10 months ago
- ☆21Dec 29, 2025Updated 8 months ago
- Inspect AI interface to Harbor tasks☆22Updated this week
- Inference API for many LLMs and other useful tools for empirical research☆136May 29, 2026Updated 3 months ago
- A python sdk for LLM finetuning and inference on runpod infrastructure☆30Updated this week
- Proteus 2.0☆10Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ⚓️ Repository for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.☆142Oct 27, 2025Updated 10 months ago
- ☆37Updated this week
- ☆18Jun 2, 2026Updated 3 months ago
- Automatic textbook formalization of Grinberg Algebraic Combinatorics☆17Jul 28, 2026Updated last month
- The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away fro…☆171Jan 20, 2026Updated 7 months ago
- One decorator for semantic caching + golden-data regression testing of LLM functions (Python).☆22Aug 20, 2026Updated 3 weeks ago
- Memory footprint reduction for transformer models☆11Jan 24, 2023Updated 3 years ago
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks