⚓️ Repository for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.
☆137Oct 27, 2025Updated 8 months ago
Alternatives and similar repositories for thought-anchors
Users that are interested in thought-anchors are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ⚓️ Interactive playground for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.☆18Dec 20, 2025Updated 7 months ago
- This was designed for interp researchers who want to do research on or with interp agents to give quality of life improvements and fix …☆146Feb 8, 2026Updated 5 months ago
- Code repo for the model organisms and convergent directions of EM papers.☆72Sep 22, 2025Updated 10 months ago
- ☆38Jul 9, 2025Updated last year
- ☆50Jul 4, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.☆78Updated this week
- ☆85Feb 18, 2026Updated 5 months ago
- Repository for the "Chain-of-Thought Reasoning In The Wild Is Not Always Faithful" paper☆35Mar 31, 2026Updated 3 months ago
- Unified access to Large Language Model modules using NNsight☆116Jul 2, 2026Updated 2 weeks ago
- ☆24Feb 13, 2026Updated 5 months ago
- ☆16Nov 14, 2025Updated 8 months ago
- Global CoT Analysis: Initial attempts to uncover patterns across many chains of thought☆20Feb 10, 2026Updated 5 months ago
- Code for Negation Neglect☆16May 22, 2026Updated 2 months ago
- The official implementation for "Mitigating Overthinking in Large Reasoning Models via Manifold Steering"☆15May 29, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University☆331Feb 8, 2026Updated 5 months ago
- Code repository for "Eliciting Secret Knowledge from Language Models"☆23Mar 30, 2026Updated 3 months ago
- ☆17Mar 16, 2026Updated 4 months ago
- A suite of interpretability tasks to evaluate agents using Scribe for notebook access☆18Oct 2, 2025Updated 9 months ago
- The nnsight package enables interpreting and manipulating the internals of deep learned models.☆995Updated this week
- Tools for exploring Transformer neuron behaviour, including input pruning and diversification.☆10Jun 6, 2023Updated 3 years ago
- ☆23Aug 30, 2025Updated 10 months ago
- Parameter Decomposition☆133Updated this week
- ☆25Jul 8, 2026Updated 2 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆15Oct 13, 2025Updated 9 months ago
- ✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks☆19Aug 16, 2024Updated last year
- ☆26Feb 20, 2026Updated 5 months ago
- ADAG: Transluce's MLP neuron-level circuit tracing library☆34Apr 10, 2026Updated 3 months ago
- ☆104Jul 15, 2026Updated last week
- ☆18Jul 9, 2025Updated last year
- A library for mechanistic interpretability of GPT-style language models☆3,705Updated this week
- Agent observability and replay tooling for AI safety & interpretability research.☆109Jun 19, 2026Updated last month
- A python sdk for LLM finetuning and inference on runpod infrastructure☆30May 12, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆25Mar 30, 2026Updated 3 months ago
- Implementations of several self-supervised pretext tasks for language and vision modalities in PyTorch.☆13Jan 19, 2021Updated 5 years ago
- ☆1,185Updated this week
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆22Dec 14, 2024Updated last year
- ☆21May 14, 2026Updated 2 months ago
- Training LLMs to Report Their Learned Behaviors☆27Apr 28, 2026Updated 2 months ago
- ☆95Apr 18, 2026Updated 3 months ago