A curated list of papers and resources on Reward Hacking, Emergent Misalignment, and Proxy Exploitation in Large Models
☆54Apr 17, 2026Updated 5 months ago
Alternatives and similar repositories for Awesome-Reward-Hacking
Users that are interested in Awesome-Reward-Hacking are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2024 main] Aligning Large Language Models with Human Preferences through Representation Engineering (https://aclanthology.org/2024.…☆28Sep 25, 2024Updated 2 years ago
- Implementation code for ACL2024:Advancing Parameter Efficiency in Fine-tuning via Representation Editing☆15Apr 20, 2024Updated 2 years ago
- ☆20Mar 17, 2025Updated last year
- [ICLR 2026] Quantile Advantage Estimation for Entropy-Safe Reasoning☆29Oct 14, 2025Updated 11 months ago
- The Official Repository for Paper "HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?"☆16May 2, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆17Jun 25, 2025Updated last year
- PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations☆14Apr 21, 2024Updated 2 years ago
- All projects I completed in Software Engineering in Fudan University from 2018 / 复旦大学2018级软件工程专业课程项目合集☆24Jan 20, 2022Updated 4 years ago
- Mobile GUI Agents under Real-world Threats: Are We There Yet?☆18Aug 4, 2026Updated last month
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆38Sep 20, 2026Updated last week
- CommerceAgentBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services☆1,268Aug 29, 2026Updated last month
- The open-source repository for PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment, which provides a general per…☆17Aug 28, 2025Updated last year
- ☆21Oct 6, 2023Updated 2 years ago
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆20Sep 14, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Jailbreak Evo☆24Jun 2, 2025Updated last year
- Companion code to https://arxiv.org/abs/2409.03797v2☆19Sep 18, 2025Updated last year
- Fast multi-modal decision model☆450Updated this week
- ☆12Jul 6, 2023Updated 3 years ago
- COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs☆18Apr 7, 2026Updated 5 months ago
- ☆173May 28, 2025Updated last year
- ☆15Apr 11, 2024Updated 2 years ago
- The official implementation of the paper "AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks"☆42Jun 1, 2026Updated 4 months ago
- LLM Safeguarding with Internal Representations☆21Apr 27, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆25May 16, 2024Updated 2 years ago
- ☆87May 8, 2026Updated 4 months ago
- HELP: a dataset for Handling Entailments with Lexical and logical Phenomena (Ver.1.0)☆15Jul 20, 2023Updated 3 years ago
- The Source Code for DR3-Eval☆40Aug 12, 2026Updated last month
- ☆16Jul 23, 2024Updated 2 years ago
- Voronoi-Based Foveated Volume Rendering☆10Sep 30, 2021Updated 5 years ago
- This repo contains code for paper: "Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach".☆26Oct 21, 2024Updated last year
- A curated list of Distribution Shift papers/articles and recent advancements.☆23Oct 20, 2022Updated 3 years ago
- ☆15Jan 24, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code and Data for "FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation" (ACL25)☆39Oct 26, 2025Updated 11 months ago
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆51Jul 6, 2026Updated 2 months ago
- ☆15Feb 10, 2026Updated 7 months ago
- ☆36Feb 27, 2025Updated last year
- ☆31Jun 9, 2026Updated 3 months ago
- ☆22May 21, 2025Updated last year
- ☆13Jun 21, 2021Updated 5 years ago