Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
☆51Jul 28, 2024Updated 2 years ago
Alternatives and similar repositories for BiPO
Users that are interested in BiPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Steering Llama 2 with Contrastive Activation Addition☆250May 23, 2024Updated 2 years ago
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- [NeurIPS 2025] Official Implementation for Learning to Steer: Input-dependent Steering for Multimodal LLMs☆19Dec 14, 2025Updated 8 months ago
- [ICLR 2025] "Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond"☆16Feb 27, 2025Updated last year
- A resource repository for representation engineering in large language models☆155Nov 14, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆21Aug 19, 2025Updated last year
- [ICLR 2025] General-purpose activation steering library☆187Sep 18, 2025Updated 11 months ago
- [WIP] [NeurIPS 2025 Spotlight] Angular Steering: Behavior Control via Rotation in Activation Space☆25May 25, 2026Updated 3 months ago
- Official codebase for "Analyzing the Generalization and Reliability of Steering Vectors"☆22Dec 14, 2024Updated last year
- ☆24Jun 13, 2024Updated 2 years ago
- ☆18Jan 17, 2024Updated 2 years ago
- [COLM 2025] SEAL: Steerable Reasoning Calibration of Large Language Models for Free☆64Apr 6, 2025Updated last year
- Camouflage poisoning via machine unlearning☆19Jul 3, 2025Updated last year
- [ACL 2025] Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms☆42Jun 4, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for "DISCO: Disentangled Communication Steering for Large Language Models" (NeurIPS 2025) https://openreview.net/pdf?id=c8AjdgdHnD☆16Oct 29, 2025Updated 10 months ago
- Code for ICML 2024 paper☆34Sep 18, 2025Updated 11 months ago
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model☆582Jan 28, 2025Updated last year
- A curated list of resources for activation engineering☆139Oct 2, 2025Updated 11 months ago
- Steering vectors for transformer language models in Pytorch / Huggingface☆163Feb 21, 2025Updated last year
- ☆18Aug 19, 2024Updated 2 years ago
- ☆18Aug 4, 2025Updated last year
- A TinyStories LM with SAEs and transcoders☆16Aug 19, 2026Updated 2 weeks ago
- Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".☆432Jun 13, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [NeurIPS 2024] "Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?"☆41Jul 18, 2025Updated last year
- [ICML25] Official repo for "Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond…☆24Sep 27, 2025Updated 11 months ago
- ☆15Mar 6, 2026Updated 5 months ago
- Code for the paper "Spectral Editing of Activations for Large Language Model Alignments"☆31Dec 20, 2024Updated last year
- [CIKM 2024] Trojan Activation Attack: Attack Large Language Models using Activation Steering for Safety-Alignment.☆30Jul 29, 2024Updated 2 years ago
- ☆71Jun 1, 2025Updated last year
- [NAACL'25 Oral] Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering☆85Jun 20, 2026Updated 2 months ago
- This is the official Gtihub repo for our paper: "BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Lang…☆24Jul 3, 2024Updated 2 years ago
- Improving Alignment and Robustness with Circuit Breakers☆268Sep 24, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆19Sep 1, 2025Updated last year
- Code for the paper "Modelling Latent Translations for Cross-Lingual Transfer"☆17Nov 22, 2021Updated 4 years ago
- [NeurIPS 2024] VeLoRA : Memory Efficient Training using Rank-1 Sub-Token Projections☆22Oct 15, 2024Updated last year
- Github Repo for ICML 2022 paper: Communication-Efficient Adaptive Federated Learning☆10Nov 18, 2022Updated 3 years ago
- [arXiv:2510.06261] "AlphaApollo: A System for Deep Agentic Reasoning"☆49Aug 21, 2026Updated 2 weeks ago
- [NeurIPS'22] Trap and Replace: Defending Backdoor Attacks by Trapping Them into an Easy-to-Replace Subnetwork. Haotao Wang, Junyuan Hong,…☆15Nov 27, 2023Updated 2 years ago
- ☆19Mar 5, 2024Updated 2 years ago