πͺ Interpreto is an interpretability toolbox for LLMs
β198Aug 28, 2026Updated this week
Alternatives and similar repositories for interpreto
Users that are interested in interpreto are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- New implementations of old orthogonal layers unlock large scale training.β33Sep 19, 2025Updated 11 months ago
- Build and train Lipschitz-constrained networks: PyTorch implementation of 1-Lipschitz layers. For TensorFlow/Keras implementation, see htβ¦β44Mar 23, 2026Updated 5 months ago
- π Influenciae is a Tensorflow Toolbox for Influence Functionsβ67Aug 24, 2026Updated last week
- π Overcomplete is a Vision-based SAE Toolboxβ150Dec 4, 2025Updated 9 months ago
- Simple, compact, and hackable post-hoc deep OOD detection for already trained tensorflow or pytorch image classifiers.β61May 19, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- β40Sep 15, 2025Updated 11 months ago
- Build and train Lipschitz constrained networks: TensorFlow implementation of k-Lipschitz layersβ102Mar 14, 2025Updated last year
- Easy-to-use MIRAGE code for faithful answer attribution in RAG applications. Paper: https://aclanthology.org/2024.emnlp-main.347/β25Mar 10, 2025Updated last year
- π Xplique is a Neural Networks Explainability Toolboxβ751Updated this week
- [NeurIPS 2023] and [ICLR 2024] for robustness certification.β10Nov 30, 2024Updated last year
- β16Nov 14, 2025Updated 9 months ago
- Unified access to Large Language Model modules using NNsightβ119Updated this week
- π Puncc is a python library for predictive uncertainty quantification using conformal prediction.β403Updated this week
- The nnsight package enables interpreting and manipulating the internals of deep learned models.β1,085Updated this week
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Generic Engine for Multi-disciplinary Scenarios, Exploration and Optimization. This is a MIRROR of our gitlab repository, the developmentβ¦β34Updated this week
- Repository for "Training Language Models To Explain Their Own Computations"β37Jul 7, 2026Updated last month
- A runway dataset and a generator of synthetic aerial images with automatic labeling.β137Updated this week
- Repository for DISRPT2021 shared taskβ16Sep 5, 2022Updated 3 years ago
- A research toolkit for decomposing and explaining text similarity across neural, structured, and symbolic levels.β31Aug 13, 2026Updated 3 weeks ago
- π¬ Interpretability for Leela Chess Zero networks.β22Updated this week
- β17Aug 19, 2026Updated 2 weeks ago
- β17May 19, 2026Updated 3 months ago
- Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profilesβ18Jul 29, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Interpretability for sequence generation models π πβ475Apr 25, 2026Updated 4 months ago
- Arrakis is a library to conduct, track and visualize mechanistic interpretability experiments.β31Jul 8, 2026Updated last month
- Code for the "Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning" paper.β18Jul 6, 2026Updated last month
- π Code for the paper: "Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity Analysis" (NeurIPS 2021)β33Jul 18, 2022Updated 4 years ago
- Code for Evaluating Explanations for Reading Comprehension with Realistic Counterfactuals.β17Apr 25, 2021Updated 5 years ago
- LENS Projectβ53Feb 22, 2024Updated 2 years ago
- ADAG: Transluce's MLP neuron-level circuit tracing libraryβ37Apr 10, 2026Updated 4 months ago
- Engine for collecting, uploading, and downloading model activationsβ31Apr 2, 2025Updated last year
- IR module for experimaestroβ14Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ICLR 2024: Energy-Based Concept Bottleneck Models: Unifying Prediction, Concept Intervention, and Probabilistic Interpretationsβ24May 1, 2025Updated last year
- Code for "On Measuring Faithfulness of Natural Language Explanations"β23Jul 14, 2026Updated last month
- Erasing conceptual knowledge from language models through low-rank fine-tuningβ23Mar 27, 2025Updated last year
- https://footprints.baulab.infoβ17Oct 4, 2024Updated last year
- Code for my NeurIPS 2024 ATTRIB paper titled "Attribution Patching Outperforms Automated Circuit Discovery"β49May 31, 2024Updated 2 years ago
- β18Oct 6, 2022Updated 3 years ago
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv Universityβ337Feb 8, 2026Updated 6 months ago