Code and materials for "Weird Generalization and Inductive Backdoors"
☆43Jan 11, 2026Updated 7 months ago
Alternatives and similar repositories for weird-generalization-and-inductive-backdoors
Users that are interested in weird-generalization-and-inductive-backdoors are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆62Jul 4, 2025Updated last year
- James' cookbook of evaluations and finetuning experiments☆35Feb 19, 2026Updated 6 months ago
- ☆16Oct 13, 2025Updated 10 months ago
- Code for Negation Neglect☆17May 22, 2026Updated 3 months ago
- Easily deploy my zsh and tmux configuration on new machines. Includes local and remote aliases to improve workflow.☆16Apr 23, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Dec 4, 2024Updated last year
- GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs☆17Nov 12, 2025Updated 9 months ago
- A collection of different ways to implement accessing and modifying internal model activations for LLMs☆24Oct 18, 2024Updated last year
- Subliminal learning in LLMs: language models can transmit hidden preferences through seemingly unrelated training data.☆25Nov 9, 2025Updated 9 months ago
- Code for the "Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning" paper.☆18Jul 6, 2026Updated 2 months ago
- ☆20Apr 10, 2025Updated last year
- ☆10May 17, 2024Updated 2 years ago
- ☆28Sep 3, 2025Updated last year
- ☆26Sep 5, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Code used to run experiments for the ICLR 2023 paper "Computational Language Acquisition with Theory of Mind".☆15Apr 27, 2023Updated 3 years ago
- ☆28Oct 6, 2024Updated last year
- ☆330Jan 12, 2026Updated 7 months ago
- ☆21Dec 10, 2025Updated 8 months ago
- ☆24Jun 22, 2021Updated 5 years ago
- A curated reading list for researchers in the Philosophy of Interpretability☆18Aug 17, 2025Updated last year
- ☆26Jan 7, 2026Updated 8 months ago
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆30Nov 20, 2024Updated last year
- Parameter Decomposition☆142Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official code release for Delta Activations: A Representation for Finetuned Large Language Models☆21Sep 5, 2025Updated last year
- A benchmark for language models based on the UK Linguistics Olympiad☆12Mar 3, 2025Updated last year
- ☆16Jun 3, 2025Updated last year
- AlgZoo: uninterpreted models with fewer than 1,500 parameters☆51Aug 19, 2026Updated 2 weeks ago
- ☆157Feb 10, 2026Updated 6 months ago
- ☆19Mar 17, 2026Updated 5 months ago
- Unified access to Large Language Model modules using NNsight☆119Updated this week
- ☆52Feb 11, 2025Updated last year
- A library for mechanistic anomaly detection☆22Jan 9, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Mar 9, 2025Updated last year
- ☆18Jul 9, 2025Updated last year
- ☆37Feb 20, 2025Updated last year
- ☆56Oct 23, 2023Updated 2 years ago
- ☆17Apr 16, 2025Updated last year
- ☆11Feb 28, 2024Updated 2 years ago
- Animal Harm Assessment public repository☆12May 3, 2026Updated 4 months ago