Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
☆50Jul 28, 2024Updated 2 years ago
Alternatives and similar repositories for BiPO
Users that are interested in BiPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Steering Llama 2 with Contrastive Activation Addition☆249May 23, 2024Updated 2 years ago
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- [ICLR 2025] "Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond"☆16Feb 27, 2025Updated last year
- A resource repository for representation engineering in large language models☆156Nov 14, 2024Updated last year
- ☆21Aug 19, 2025Updated 11 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2025] General-purpose activation steering library☆181Sep 18, 2025Updated 10 months ago
- [WIP] [NeurIPS 2025 Spotlight] Angular Steering: Behavior Control via Rotation in Activation Space☆25May 25, 2026Updated 2 months ago
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆22Dec 14, 2024Updated last year
- ☆18Jan 17, 2024Updated 2 years ago
- [COLM 2025] SEAL: Steerable Reasoning Calibration of Large Language Models for Free☆62Apr 6, 2025Updated last year
- Camouflage poisoning via machine unlearning☆19Jul 3, 2025Updated last year
- [ACL 2025] Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms☆42Jun 4, 2025Updated last year
- Code for "DISCO: Disentangled Communication Steering for Large Language Models" (NeurIPS 2025) https://openreview.net/pdf?id=c8AjdgdHnD☆16Oct 29, 2025Updated 9 months ago
- Code for ICML 2024 paper☆34Sep 18, 2025Updated 10 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model☆581Jan 28, 2025Updated last year
- A curated list of resources for activation engineering☆139Oct 2, 2025Updated 10 months ago
- ☆18Aug 19, 2024Updated last year
- ☆419Jul 2, 2024Updated 2 years ago
- ☆19Aug 4, 2025Updated last year
- Evaluate interpretability methods on localizing and disentangling concepts in LLMs.☆58Oct 30, 2025Updated 9 months ago
- A TinyStories LM with SAEs and transcoders☆15Apr 3, 2025Updated last year
- Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".☆428Jun 13, 2025Updated last year
- [NeurIPS 2024] "Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?"☆41Jul 18, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICML25] Official repo for "Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond…☆24Sep 27, 2025Updated 10 months ago
- ☆15Mar 6, 2026Updated 5 months ago
- Stanford NLP Python library for benchmarking the utility of LLM interpretability methods☆211Mar 12, 2026Updated 5 months ago
- [CIKM 2024] Trojan Activation Attack: Attack Large Language Models using Activation Steering for Safety-Alignment.☆30Jul 29, 2024Updated 2 years ago
- ☆71Jun 1, 2025Updated last year
- [NAACL'25 Oral] Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering☆83Jun 20, 2026Updated last month
- Improving Alignment and Robustness with Circuit Breakers☆267Sep 24, 2024Updated last year
- ☆19Sep 1, 2025Updated 11 months ago
- Code for the paper "Modelling Latent Translations for Cross-Lingual Transfer"☆17Nov 22, 2021Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [NeurIPS 2024] VeLoRA : Memory Efficient Training using Rank-1 Sub-Token Projections☆22Oct 15, 2024Updated last year
- ☆24Nov 11, 2024Updated last year
- Github Repo for ICML 2022 paper: Communication-Efficient Adaptive Federated Learning☆10Nov 18, 2022Updated 3 years ago
- [arXiv:2510.06261] "AlphaApollo: A System for Deep Agentic Reasoning"☆49May 18, 2026Updated 2 months ago
- [NeurIPS'22] Trap and Replace: Defending Backdoor Attacks by Trapping Them into an Easy-to-Replace Subnetwork. Haotao Wang, Junyuan Hong,…☆15Nov 27, 2023Updated 2 years ago
- ☆28Mar 4, 2025Updated last year
- Code for In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering☆202Feb 13, 2025Updated last year