Real-Time Detection of Hallucinated Entities in Long-Form Generation
☆295Nov 16, 2025Updated 10 months ago
Alternatives and similar repositories for hallucination_probes
Users that are interested in hallucination_probes are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A suite of interpretability tasks to evaluate agents using Scribe for notebook access☆18Oct 2, 2025Updated 11 months ago
- ☆87Feb 18, 2026Updated 7 months ago
- ☆29Nov 6, 2025Updated 10 months ago
- A toolkit for embedding text datasets with sparse autoencoders☆31Mar 24, 2026Updated 6 months ago
- ☆20Mar 16, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Parameter Decomposition☆143Sep 15, 2026Updated last week
- Open source interpretability artefacts for R1.☆183Apr 21, 2025Updated last year
- ☆19Feb 12, 2026Updated 7 months ago
- Unified access to Large Language Model modules using NNsight☆120Updated this week
- [NAACL'25 Oral] Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering☆85Jun 20, 2026Updated 3 months ago
- ☆19Jun 13, 2023Updated 3 years ago
- A collection of lightweight interpretability scripts to understand how LLMs think☆92Mar 18, 2026Updated 6 months ago
- ☆18Nov 5, 2025Updated 10 months ago
- A python sdk for LLM finetuning and inference on runpod infrastructure☆30Sep 7, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆20Apr 10, 2025Updated last year
- Engine for collecting, uploading, and downloading model activations☆31Apr 2, 2025Updated last year
- Repo for Paper: Discovering Interpretable Algorithms by Decompiling Transformers to RASP☆16May 25, 2026Updated 3 months ago
- This was designed for interp researchers who want to do research on or with interp agents to give quality of life improvements and fix …☆146Feb 8, 2026Updated 7 months ago
- ☆79Mar 6, 2025Updated last year
- ☆2,911Updated this week
- ☆15Feb 21, 2024Updated 2 years ago
- ☆12Jan 10, 2023Updated 3 years ago
- Training LLMs to Report Their Learned Behaviors☆29Apr 28, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- https://arxiv.org/abs/2404.10917☆14Mar 18, 2025Updated last year
- A toolkit for describing model features and intervening on those features to steer behavior.☆262Mar 16, 2026Updated 6 months ago
- The nnsight package enables interpreting and manipulating the internals of deep learned models.☆1,108Updated this week
- Inference API for many LLMs and other useful tools for empirical research☆137May 29, 2026Updated 3 months ago
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models☆464Apr 22, 2026Updated 5 months ago
- Code for Negation Neglect☆18May 22, 2026Updated 4 months ago
- An implementation of "Subspace Representations for Soft Set Operations and Sentence Similarities" (NAACL 2024)☆10May 31, 2024Updated 2 years ago
- Training Sparse Autoencoders on Language Models☆1,540Updated this week
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.☆86Sep 1, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Delphi was the home of a temple to Phoebus Apollo, which famously had the inscription, 'Know Thyself.' This library lets language models …☆275Updated this week
- Pipeline parallel training on Apple Silicon!☆34Dec 5, 2025Updated 9 months ago
- ☆1,318Updated this week
- Subliminal learning in LLMs: language models can transmit hidden preferences through seemingly unrelated training data.☆25Nov 9, 2025Updated 10 months ago
- Accelerate common Petri dish assays with AI.☆14Oct 28, 2025Updated 10 months ago
- ☆64Jul 4, 2025Updated last year
- Repository for the paper: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning☆18Feb 21, 2025Updated last year