Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University
☆333Feb 8, 2026Updated 5 months ago
Alternatives and similar repositories for llm-interp-tau
Users that are interested in llm-interp-tau are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This was designed for interp researchers who want to do research on or with interp agents to give quality of life improvements and fix …☆146Feb 8, 2026Updated 5 months ago
- ☆16Nov 14, 2025Updated 8 months ago
- Repository for "Training Language Models To Explain Their Own Computations"☆23Jul 7, 2026Updated 3 weeks ago
- ADAG: Transluce's MLP neuron-level circuit tracing library☆34Apr 10, 2026Updated 3 months ago
- ⚓️ Repository for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.☆137Oct 27, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Training Sparse Autoencoders on Language Models☆1,487Updated this week
- The nnsight package enables interpreting and manipulating the internals of deep learned models.☆1,005Updated this week
- Extract residual-stream activations and apply steering vectors (including activation oracles) to any vLLM model during inference.☆117Updated this week
- Unified access to Large Language Model modules using NNsight☆116Updated this week
- A library for mechanistic interpretability of GPT-style language models☆3,723Updated this week
- Code repo for the model organisms and convergent directions of EM papers.☆72Sep 22, 2025Updated 10 months ago
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.☆78Jul 20, 2026Updated last week
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆20Jan 19, 2025Updated last year
- ☆50Jul 4, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Code repository for "Eliciting Secret Knowledge from Language Models"☆24Mar 30, 2026Updated 3 months ago
- ☆86Feb 25, 2025Updated last year
- Engine for collecting, uploading, and downloading model activations☆30Apr 2, 2025Updated last year
- An LLM agent framework for automated AI interpretability research☆18Apr 17, 2026Updated 3 months ago
- This repository includes code for the paper "Does Localization Inform Editing? Surprising Differences in Where Knowledge Is Stored vs. Ca…☆62May 9, 2023Updated 3 years ago
- ☆16Jul 7, 2026Updated 3 weeks ago
- How do transformer LMs encode relations?☆60Feb 24, 2024Updated 2 years ago
- A curated reading list for researchers in the Philosophy of Interpretability☆17Aug 17, 2025Updated 11 months ago
- ☆1,192Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- 🪝PISCES - Precise In-Parameter Suppression for Concept EraSure in Large Language Models☆13Jun 28, 2026Updated last month
- Implementations of several self-supervised pretext tasks for language and vision modalities in PyTorch.☆13Jan 19, 2021Updated 5 years ago
- A library for training crosscoders☆17May 28, 2025Updated last year
- ☆15Oct 17, 2023Updated 2 years ago
- Parameter Decomposition☆136Updated this week
- The Github repo for our survey paper: "Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large…☆150Apr 15, 2026Updated 3 months ago
- ☆18Jul 9, 2025Updated last year
- Code for Negation Neglect☆16May 22, 2026Updated 2 months ago
- Sparsify transformers with SAEs and transcoders☆735Jul 20, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A neural RST discourse parser with well pre-trained XLNet.☆17Jun 13, 2022Updated 4 years ago
- awesome SAE papers☆79May 24, 2025Updated last year
- Code for the "Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning" paper.☆17Jul 6, 2026Updated 3 weeks ago
- ☆25Jul 8, 2026Updated 3 weeks ago
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆269Updated this week
- Attribution-based Parameter Decomposition☆35Jun 11, 2025Updated last year
- ☆2,879Jul 18, 2026Updated last week