Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
☆51Jul 28, 2024Updated 2 years ago
Alternatives and similar repositories for BiPO
Users that are interested in BiPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Steering Llama 2 with Contrastive Activation Addition☆253May 23, 2024Updated 2 years ago
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- [NeurIPS 2025] Official Implementation for Learning to Steer: Input-dependent Steering for Multimodal LLMs☆20Dec 14, 2025Updated 9 months ago
- [ICLR 2025] "Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond"☆16Feb 27, 2025Updated last year
- A resource repository for representation engineering in large language models☆156Nov 14, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆23Aug 19, 2025Updated last year
- [ICLR 2025] General-purpose activation steering library☆190Sep 18, 2025Updated last year
- [WIP] [NeurIPS 2025 Spotlight] Angular Steering: Behavior Control via Rotation in Activation Space☆25May 25, 2026Updated 3 months ago
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆23Dec 14, 2024Updated last year
- ☆24Jun 13, 2024Updated 2 years ago
- ☆18Jan 17, 2024Updated 2 years ago
- [COLM 2025] SEAL: Steerable Reasoning Calibration of Large Language Models for Free☆65Apr 6, 2025Updated last year
- [ACL 2025] Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms☆42Jun 4, 2025Updated last year
- Code for "DISCO: Disentangled Communication Steering for Large Language Models" (NeurIPS 2025) https://openreview.net/pdf?id=c8AjdgdHnD☆16Oct 29, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for ICML 2024 paper☆34Sep 18, 2025Updated last year
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model☆587Jan 28, 2025Updated last year
- A curated list of resources for activation engineering☆141Oct 2, 2025Updated 11 months ago
- Steering vectors for transformer language models in Pytorch / Huggingface☆163Feb 21, 2025Updated last year
- ☆430Jul 2, 2024Updated 2 years ago
- ☆18Aug 4, 2025Updated last year
- Unlearnable Examples Give a False Sense of Security: Piercing through Unexploitable Data with Learnable Examples☆11Oct 14, 2024Updated last year
- Evaluate interpretability methods on localizing and disentangling concepts in LLMs.☆58Oct 30, 2025Updated 10 months ago
- A TinyStories LM with SAEs and transcoders☆16Aug 19, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".☆446Jun 13, 2025Updated last year
- [NeurIPS 2024] "Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?"☆41Jul 18, 2025Updated last year
- [ICML25] Official repo for "Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond…☆25Sep 27, 2025Updated 11 months ago
- ☆15Mar 6, 2026Updated 6 months ago
- Stanford NLP Python library for benchmarking the utility of LLM interpretability methods☆216Mar 12, 2026Updated 6 months ago
- [CIKM 2024] Trojan Activation Attack: Attack Large Language Models using Activation Steering for Safety-Alignment.☆30Jul 29, 2024Updated 2 years ago
- ☆71Jun 1, 2025Updated last year
- [NAACL'25 Oral] Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering☆85Jun 20, 2026Updated 3 months ago
- PyTorch implementation of our ICLR 2023 paper titled "Is Adversarial Training Really a Silver Bullet for Mitigating Data Poisoning?".☆12Mar 13, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Improving Alignment and Robustness with Circuit Breakers☆269Sep 24, 2024Updated 2 years ago
- Awesome Multimodal Fusion in Speech Emotion Recognition☆17Nov 11, 2025Updated 10 months ago
- ☆19Sep 1, 2025Updated last year
- Code for the paper "Modelling Latent Translations for Cross-Lingual Transfer"☆17Nov 22, 2021Updated 4 years ago
- [NeurIPS 2024] VeLoRA : Memory Efficient Training using Rank-1 Sub-Token Projections☆22Oct 15, 2024Updated last year
- ☆24Nov 11, 2024Updated last year
- Github Repo for ICML 2022 paper: Communication-Efficient Adaptive Federated Learning☆10Nov 18, 2022Updated 3 years ago