☆16Mar 13, 2025Updated last year
Alternatives and similar repositories for ConceptorSteering
Users that are interested in ConceptorSteering are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A very hacky set of functions for getting plotly to do what I want when doing mech interp research, designed to be compatible with PyTorc…☆15Jun 16, 2023Updated 3 years ago
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆12Oct 31, 2024Updated last year
- Interactive visualization tool for multiple volumes, meshes and points based on VTK. The app can be controlled with Python scripts (optio…☆11Mar 13, 2022Updated 4 years ago
- Official code for Conformal Isometry of Lie Group Representation in Recurrent Network of Grid Cells (NeurIPS workshop on Symmetry and Geo…☆13Nov 1, 2022Updated 3 years ago
- ☆19Sep 1, 2025Updated 11 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A tiny easily hackable implementation of a feature dashboard.☆17Oct 21, 2025Updated 9 months ago
- Trains Sparse Autoencoders based on outputs from language models☆11Oct 7, 2024Updated last year
- Tools for exploring Transformer neuron behaviour, including input pruning and diversification.☆24Sep 28, 2023Updated 2 years ago
- ☆12Jan 10, 2025Updated last year
- A benchmark for mechanistic discovery of circuits in Transformers☆18Dec 15, 2024Updated last year
- A Mechanistic Interpretability Analysis of Grokking☆29Sep 26, 2022Updated 3 years ago
- This repository generates multi-view observations for static/dynamic CLEVR scenes.☆12Apr 17, 2023Updated 3 years ago
- Latent Space Geometry for Neural Networks in Python☆19Jun 1, 2026Updated 2 months ago
- ☆18Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A library of visualization tools for the interpretability and hallucination analysis of large vision-language models (LVLMs).☆43May 22, 2025Updated last year
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆29Nov 20, 2024Updated last year
- ☆42May 21, 2025Updated last year
- An exploration of LLM steering☆28Jun 15, 2024Updated 2 years ago
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆22Dec 14, 2024Updated last year
- [ICLR26] Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs☆24Apr 8, 2026Updated 4 months ago
- Code for NeurIPS 2022 No Free Lunch from Deep Learning in Neuroscience: A Case Study through Models of the Entorhinal-Hippocampal Circuit☆15Nov 11, 2022Updated 3 years ago
- Library for running trials with parameter variation☆13Oct 14, 2018Updated 7 years ago
- Flexible modulation of sequence generation in the entorhinal-hippocampal system.☆13Apr 5, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Genertaes control vectors for use with llama.cpp in GGUF format.☆48Mar 19, 2025Updated last year
- Steering Llama 2 with Contrastive Activation Addition☆248May 23, 2024Updated 2 years ago
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- Lab tutorials for the MSc NLP course at the University of Groningen 🐮☆29Feb 25, 2025Updated last year
- ☆41Jul 3, 2026Updated last month
- ☆34Nov 16, 2025Updated 8 months ago
- Building Llama 3 from scratch using PyTorch☆13Sep 1, 2024Updated last year
- Draft of the SPL Multi-modal brain atlas☆19Dec 12, 2017Updated 8 years ago
- ☆33Nov 28, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Evaluate interpretability methods on localizing and disentangling concepts in LLMs.☆58Oct 30, 2025Updated 9 months ago
- ☆21Aug 19, 2025Updated 11 months ago
- Accompanying codebase for neuroscope.io, a website for displaying max activating dataset examples for language model neurons☆15Feb 13, 2023Updated 3 years ago
- Multi-Layer Sparse Autoencoders (ICLR 2025)☆30Feb 6, 2026Updated 6 months ago
- Tools for optimizing steering vectors in LLMs.☆22Apr 10, 2025Updated last year
- ☆18Jun 19, 2023Updated 3 years ago
- Neurorobotics 2020☆19Aug 14, 2023Updated 2 years ago