This repository contains the code and data for the paper "SelfIE: Self-Interpretation of Large Language Model Embeddings" by Haozhe Chen, Carl Vondrick, and Chengzhi Mao.
☆59Dec 9, 2024Updated last year
Alternatives and similar repositories for selfie
Users that are interested in selfie are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repository contains the code used for the experiments in the paper "Language Models use Lookbacks to Track Beliefs".☆17Mar 14, 2026Updated 6 months ago
- ☆26Dec 20, 2023Updated 2 years ago
- ☆36Nov 16, 2025Updated 10 months ago
- Playing around with various jailbreaking techniques ahead of the Gray Swan AI Ultimate Jailbreaking Competition☆19Oct 6, 2024Updated last year
- All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks☆17Apr 24, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- code of paper "Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM"☆14Nov 17, 2023Updated 2 years ago
- ☆12Oct 23, 2022Updated 3 years ago
- ☆30Aug 2, 2024Updated 2 years ago
- Learning from preferences is a common paradigm for fine-tuning language models. Yet, many algorithmic design decisions come into play. Ou…☆32Apr 20, 2024Updated 2 years ago
- Providing the answer to "How to do patching on all available SAEs on GPT-2?". It is an official repository of the implementation of the p…☆14Jan 26, 2025Updated last year
- Create feature-centric and prompt-centric visualizations for sparse autoencoders (like those from Anthropic's published research).☆275Feb 27, 2026Updated 6 months ago
- TACL 2025: Investigating Adversarial Trigger Transfer in Large Language Models☆20Aug 17, 2025Updated last year
- ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions☆15Jun 28, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆16Jan 16, 2025Updated last year
- ☆187May 1, 2026Updated 4 months ago
- ☆17Updated this week
- Social Network Analysis and STEM Education is designed to prepare researchers to apply network analysis in order to better understand and…☆15Jul 14, 2025Updated last year
- This is the oficial repository for "Safer-Instruct: Aligning Language Models with Automated Preference Data"☆17Feb 22, 2024Updated 2 years ago
- [EMNLP 2026] PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization☆18Aug 21, 2026Updated last month
- [ICLR 2025] Code&Data for the paper "Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization"☆15Jun 21, 2024Updated 2 years ago
- ☆12Apr 25, 2025Updated last year
- Using sparse coding to find distributed representations used by neural networks.☆313Nov 10, 2023Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals; ACL 2024☆13May 24, 2024Updated 2 years ago
- ☆14Mar 11, 2024Updated 2 years ago
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆23Dec 14, 2024Updated last year
- Code for the NAACL 2024 HCI+NLP Workshop paper "LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tool…☆12Mar 24, 2024Updated 2 years ago
- ☆17Aug 1, 2025Updated last year
- ☆33Nov 28, 2024Updated last year
- ☆16Jul 23, 2024Updated 2 years ago
- ☆35Nov 7, 2024Updated last year
- Enhancing Large Vision Language Models with Self-Training on Image Comprehension.☆68May 31, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code release for "Debating with More Persuasive LLMs Leads to More Truthful Answers"☆132Mar 22, 2024Updated 2 years ago
- EMNLP 2024: Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue☆37May 26, 2025Updated last year
- This is the official code for the paper "Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable".☆35Mar 11, 2025Updated last year
- Interpretating the latent space representations of attention head outputs for LLMs☆39Aug 13, 2024Updated 2 years ago
- WMDP is a LLM proxy benchmark for hazardous knowledge in bio, cyber, and chemical security. We also release code for RMU, an unlearning m…☆184May 29, 2025Updated last year
- ☆225Oct 14, 2025Updated 11 months ago
- ☆18Jun 3, 2026Updated 3 months ago