Repository for the "Chain-of-Thought Reasoning In The Wild Is Not Always Faithful" paper
☆35Mar 31, 2026Updated 3 months ago
Alternatives and similar repositories for chainscope
Users that are interested in chainscope are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆57Oct 23, 2023Updated 2 years ago
- Improving transparency of large language models' reasoning☆15Nov 25, 2025Updated 8 months ago
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆22Dec 14, 2024Updated last year
- ☆95Oct 8, 2025Updated 9 months ago
- ⚓️ Repository for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.☆137Oct 27, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official Code for our paper: "Language Models Learn to Mislead Humans via RLHF""☆20Oct 11, 2024Updated last year
- Code for Negation Neglect☆16May 22, 2026Updated 2 months ago
- Open Source Replication of Anthropic's Alignment Faking Paper☆58Apr 4, 2025Updated last year
- ☆17Mar 16, 2026Updated 4 months ago
- ☆16Nov 14, 2025Updated 8 months ago
- ☆39Jul 9, 2025Updated last year
- ☆141Oct 28, 2023Updated 2 years ago
- Implementation of the Decrypto benchmark for multi-agent reasoning and theory of mind.☆22Jan 19, 2026Updated 6 months ago
- Code for Preventing Language Models From Hiding Their Reasoning, which evaluates defenses against LLM steganography.☆25Jan 26, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆12Jun 17, 2024Updated 2 years ago
- ☆15Jun 17, 2025Updated last year
- Code for the EACL 2024 paper: "Small Language Models Improve Giants by Rewriting Their Outputs"☆12Apr 20, 2024Updated 2 years ago
- Experimental LLM interface exploring new ways to use AI to improve human thinking☆21Apr 13, 2026Updated 3 months ago
- A benchmark for language models based on the UK Linguistics Olympiad☆12Mar 3, 2025Updated last year
- ☆24Feb 13, 2026Updated 5 months ago
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- Representational similarity consistency across dataset and their driving factors.☆16Jun 20, 2025Updated last year
- ☆114Aug 8, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Code for "On Measuring Faithfulness of Natural Language Explanations"☆23Jul 14, 2026Updated 2 weeks ago
- ☆25May 25, 2024Updated 2 years ago
- Creating a game to play Figgie & Train an agent to play against☆15Dec 3, 2022Updated 3 years ago
- Persistent caching for Python functions☆19Dec 10, 2025Updated 7 months ago
- [NeurIPS 2025 MechInterp Workshop - Spotlight] Official implementation of the paper "RelP: Faithful and Efficient Circuit Discovery in La…