Official code for Steering Large Language Models using Conceptors, presented at the NeurIPS 2024 MINT Workshop.
☆16Mar 13, 2025Updated last year
Alternatives and similar repositories for ConceptorSteering
Users that are interested in ConceptorSteering are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A very hacky set of functions for getting plotly to do what I want when doing mech interp research, designed to be compatible with PyTorc…☆15Jun 16, 2023Updated 3 years ago
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆11Oct 31, 2024Updated last year
- ☆19Sep 1, 2025Updated last year
- A tiny easily hackable implementation of a feature dashboard.☆18Oct 21, 2025Updated 11 months ago
- Trains Sparse Autoencoders based on outputs from language models☆11Oct 7, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆19Mar 5, 2024Updated 2 years ago
- ☆12Jan 10, 2025Updated last year
- A benchmark for mechanistic discovery of circuits in Transformers☆18Dec 15, 2024Updated last year
- A Mechanistic Interpretability Analysis of Grokking☆29Sep 26, 2022Updated 3 years ago
- ☆17Updated this week
- A library of visualization tools for the interpretability and hallucination analysis of large vision-language models (LVLMs).☆43May 22, 2025Updated last year
- Improving Steering Vectors by Targeting Sparse Autoencoder Features☆30Nov 20, 2024Updated last year
- ☆42May 21, 2025Updated last year
- An exploration of LLM steering☆28Jun 15, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆23Dec 14, 2024Updated last year
- Genertaes control vectors for use with llama.cpp in GGUF format.☆49Mar 19, 2025Updated last year
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- Steering Llama 2 with Contrastive Activation Addition☆251May 23, 2024Updated 2 years ago
- [ICLR26] Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs☆25Apr 8, 2026Updated 5 months ago
- ☆43Jul 3, 2026Updated 2 months ago
- ☆36Nov 16, 2025Updated 10 months ago
- Experiments with representation engineering☆14Feb 28, 2024Updated 2 years ago
- Building Llama 3 from scratch using PyTorch☆13Sep 1, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆33Nov 28, 2024Updated last year
- Codebase for Obfuscated Activations Bypass LLM Latent-Space Defenses☆33Feb 11, 2025Updated last year
- Evaluate interpretability methods on localizing and disentangling concepts in LLMs.☆58Oct 30, 2025Updated 10 months ago
- ☆22Aug 19, 2025Updated last year
- Accompanying codebase for neuroscope.io, a website for displaying max activating dataset examples for language model neurons☆15Feb 13, 2023Updated 3 years ago
- An API that detect expiration date from the product package's picture based on Deep Learning Algorithms☆11Jun 4, 2022Updated 4 years ago
- Multi-Layer Sparse Autoencoders (ICLR 2025)☆30Feb 6, 2026Updated 7 months ago
- ☆18Jun 19, 2023Updated 3 years ago
- ☆12Aug 31, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [TMLR 25] An automated method for explaining complex neuron behaviors in deep vision models using large language models☆12Feb 20, 2025Updated last year
- ☆17Jul 4, 2025Updated last year
- Codes and data for AAAI-24 paper "Advancing Spatial Reasoning in Large Language Models: An In-depth Evaluation and Enhancement Using the …☆14Apr 23, 2024Updated 2 years ago
- ☆17Nov 5, 2024Updated last year
- ☆17Oct 29, 2025Updated 10 months ago
- Erasing conceptual knowledge from language models through low-rank fine-tuning☆25Mar 27, 2025Updated last year
- ASIDE: Architectural Separation of Instructions and Data in Language Models [ICLR 2026]☆18Jun 10, 2026Updated 3 months ago