A curated list of papers and resources on Reward Hacking, Emergent Misalignment, and Proxy Exploitation in Large Models
☆43Apr 17, 2026Updated 3 months ago
Alternatives and similar repositories for Awesome-Reward-Hacking
Users that are interested in Awesome-Reward-Hacking are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2024 main] Aligning Large Language Models with Human Preferences through Representation Engineering (https://aclanthology.org/2024.…☆28Sep 25, 2024Updated last year
- Implementation code for ACL2024:Advancing Parameter Efficiency in Fine-tuning via Representation Editing☆15Apr 20, 2024Updated 2 years ago
- ☆20Mar 17, 2025Updated last year
- [ICLR 2026] Quantile Advantage Estimation for Entropy-Safe Reasoning☆29Oct 14, 2025Updated 9 months ago
- The Official Repository for Paper "HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?"☆15May 2, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆17Jun 25, 2025Updated last year
- Official pytorch implementation of Spatial Relation Decomposition method (AAAI 23)☆23Dec 29, 2023Updated 2 years ago
- All projects I completed in Software Engineering in Fudan University from 2018 / 复旦大学2018级软件工程专业课程项目合集☆22Jan 20, 2022Updated 4 years ago
- Mobile GUI Agents under Real-world Threats: Are We There Yet?☆18May 18, 2026Updated 2 months ago
- Accompanying code for our NeurIPS 2019 paper☆11Nov 7, 2019Updated 6 years ago
- MultiPriv offers multilingual, multimodal PII entities and prompts for studying privacy risks in LLMs/VLMs. It also supports broader PII-…☆36Jun 29, 2026Updated last month
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 5 months ago
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆18Jul 1, 2026Updated last month
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆11Dec 2, 2018Updated 7 years ago
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆21Apr 24, 2026Updated 3 months ago
- The official implementation of the paper "AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks"☆29Jun 1, 2026Updated 2 months ago
- ☆170May 28, 2025Updated last year
- LLM Safeguarding with Internal Representations☆20Apr 27, 2026Updated 3 months ago
- ☆77May 8, 2026Updated 2 months ago
- ☆25May 16, 2024Updated 2 years ago
- HELP: a dataset for Handling Entailments with Lexical and logical Phenomena (Ver.1.0)☆15Jul 20, 2023Updated 3 years ago
- ☆39May 7, 2026Updated 2 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Voronoi-Based Foveated Volume Rendering☆10Sep 30, 2021Updated 4 years ago
- A Security Benchmark for Claude Code Agent Skills☆69Updated this week
- This repo contains code for paper: "Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach".☆26Oct 21, 2024Updated last year
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆24Apr 24, 2026Updated 3 months ago
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆46Jul 6, 2026Updated 3 weeks ago
- Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs☆45Jun 17, 2025Updated last year
- ☆15Jan 24, 2025Updated last year
- Official repository for Beyond Binary Rewards: Training LMs to Reason about Their Uncertainty☆68Aug 20, 2025Updated 11 months ago
- (ACL-2025 main conference) Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback☆44Jun 24, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code and Data for "FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation" (ACL25)☆39Oct 26, 2025Updated 9 months ago
- ☆51Jun 14, 2024Updated 2 years ago
- 🚀 Sliding Window Attention Training for Efficient Large Language Models☆19Jun 7, 2026Updated last month
- Shared repository for TRIP dataset for verifiable NLU and coherence measurement for text classifiers.☆16Nov 15, 2022Updated 3 years ago
- Arabic To English translation using transformer neural nets.☆15Mar 15, 2019Updated 7 years ago
- ☆15Feb 10, 2026Updated 5 months ago
- ☆33Feb 27, 2025Updated last year