A curated list of papers and resources on Reward Hacking, Emergent Misalignment, and Proxy Exploitation in Large Models
☆48Apr 17, 2026Updated 4 months ago
Alternatives and similar repositories for Awesome-Reward-Hacking
Users that are interested in Awesome-Reward-Hacking are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2024 main] Aligning Large Language Models with Human Preferences through Representation Engineering (https://aclanthology.org/2024.…☆28Sep 25, 2024Updated last year
- Implementation code for ACL2024:Advancing Parameter Efficiency in Fine-tuning via Representation Editing☆15Apr 20, 2024Updated 2 years ago
- ☆20Mar 17, 2025Updated last year
- [ICLR 2026] Quantile Advantage Estimation for Entropy-Safe Reasoning☆29Oct 14, 2025Updated 10 months ago
- The Official Repository for Paper "HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?"☆16May 2, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆17Jun 25, 2025Updated last year
- Official pytorch implementation of Spatial Relation Decomposition method (AAAI 23)☆23Dec 29, 2023Updated 2 years ago
- PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations☆13Apr 21, 2024Updated 2 years ago
- Mobile GUI Agents under Real-world Threats: Are We There Yet?☆18Aug 4, 2026Updated 2 weeks ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 5 months ago
- The open-source repository for PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment, which provides a general per…☆17Aug 28, 2025Updated 11 months ago
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆18Jul 1, 2026Updated last month
- Jailbreak Evo☆23Jun 2, 2025Updated last year
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Companion code to https://arxiv.org/abs/2409.03797v2☆19Sep 18, 2025Updated 11 months ago
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆22Apr 24, 2026Updated 3 months ago
- COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs☆18Apr 7, 2026Updated 4 months ago
- The official implementation of the paper "AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks"☆30Jun 1, 2026Updated 2 months ago
- LLM Safeguarding with Internal Representations☆20Apr 27, 2026Updated 3 months ago
- ☆80May 8, 2026Updated 3 months ago
- ☆25May 16, 2024Updated 2 years ago
- [ACL'23 Findings] This is the code repo for our ACL'23 Findings paper "ReGen: Zero-Shot Text Classification via Training Data Generation …☆24Sep 8, 2023Updated 2 years ago
- The Source Code for DR3-Eval☆39Aug 12, 2026Updated last week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [COLM 2026] Official implementation for "MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo…☆20Apr 23, 2026Updated 4 months ago
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆24Apr 24, 2026Updated 4 months ago
- A curated list of Distribution Shift papers/articles and recent advancements.☆23Oct 20, 2022Updated 3 years ago
- Interpreting Learned Search and Planning: Reverse-engineering recurrent convolutional networks (DRC) that play Sokoban☆22Jun 29, 2025Updated last year
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆47Jul 6, 2026Updated last month
- Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs☆45Jun 17, 2025Updated last year
- ☆15Jan 24, 2025Updated last year
- Official repository for Beyond Binary Rewards: Training LMs to Reason about Their Uncertainty☆69Aug 20, 2025Updated last year
- (ACL-2025 main conference) Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback☆44Jun 24, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Code and Data for "FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation" (ACL25)☆39Oct 26, 2025Updated 9 months ago
- 🚀 Sliding Window Attention Training for Efficient Large Language Models☆20Jun 7, 2026Updated 2 months ago
- ☆15Feb 10, 2026Updated 6 months ago
- ☆34Feb 27, 2025Updated last year
- ☆29Jun 9, 2026Updated 2 months ago
- [NDSS 2026] Official repo for Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography☆60Mar 14, 2026Updated 5 months ago
- [COLM '24] Source-Aware Training Enables Knowledge Attribution in Language Models☆20Apr 1, 2025Updated last year