WMDP is a LLM proxy benchmark for hazardous knowledge in bio, cyber, and chemical security. We also release code for RMU, an unlearning method which reduces LLM performance on WMDP while retaining general capabilities.
☆177May 29, 2025Updated last year
Alternatives and similar repositories for wmdp
Users that are interested in wmdp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆28Oct 6, 2024Updated last year
- ☆32Aug 9, 2024Updated 2 years ago
- [NeurIPS D&B '25] The one-stop repository for LLM unlearning☆586Mar 18, 2026Updated 5 months ago
- Official repo for EMNLP'24 paper "SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning"☆30Oct 1, 2024Updated last year
- [NeurIPS25] Official repo for "Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning"☆46Oct 3, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR 2025] A Closer Look at Machine Unlearning for Large Language Models☆49Dec 4, 2024Updated last year
- Official repo for NeurIPS'24 paper "WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models"☆19Dec 16, 2024Updated last year
- Improving Alignment and Robustness with Circuit Breakers☆267Sep 24, 2024Updated last year
- ☆34Mar 13, 2025Updated last year
- RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models. NeurIPS 2024☆100Sep 30, 2024Updated last year
- LLM Unlearning☆188Oct 20, 2023Updated 2 years ago
- ☆48Oct 1, 2024Updated last year
- A resource repository for machine unlearning in large language models☆623Aug 6, 2026Updated last week
- [NeurIPS 2024] Large Language Model Unlearning via Embedding-Corrupted Prompts☆41Sep 26, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2025] Official Repository for "Tamper-Resistant Safeguards for Open-Weight LLMs"☆70Jun 9, 2025Updated last year
- ☆76Jul 15, 2024Updated 2 years ago
- ☆18Oct 12, 2025Updated 10 months ago
- [ICLR 2025] FLAT: LLM Unlearning via Loss Adjustment with Only Forget Data☆14Feb 26, 2025Updated last year
- [ICML25] Official repo for "Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond…☆24Sep 27, 2025Updated 10 months ago
- "In-Context Unlearning: Language Models as Few Shot Unlearners". Martin Pawelczyk, Seth Neel* and Himabindu Lakkaraju*; ICML 2024.☆31Oct 18, 2023Updated 2 years ago
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal☆1,028Aug 16, 2024Updated 2 years ago
- ☆188Apr 22, 2026Updated 3 months ago
- Official Implementation of "Learning to Refuse: Towards Mitigating Privacy Risks in LLMs"☆10Dec 13, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ACL 2025] Knowledge Unlearning for Large Language Models☆49Sep 18, 2025Updated 11 months ago
- ☆19Jun 21, 2025Updated last year
- This is the official code for the paper "Vaccine: Perturbation-aware Alignment for Large Language Models" (NeurIPS2024)☆52Jan 15, 2026Updated 7 months ago
- ☆48Sep 29, 2024Updated last year
- ☆154Jul 23, 2025Updated last year
- We jailbreak GPT-3.5 Turbo’s safety guardrails by fine-tuning it on only 10 adversarially designed examples, at a cost of less than $0.20…☆357Feb 23, 2024Updated 2 years ago
- [NeurIPS 2024 D&B] Evaluating Copyright Takedown Methods for Language Models☆17Jul 17, 2024Updated 2 years ago
- Steering Llama 2 with Contrastive Activation Addition☆249May 23, 2024Updated 2 years ago
- ☆46Feb 11, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆29Feb 25, 2025Updated last year
- Röttger et al. (NAACL 2024): "XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models"☆142Feb 24, 2025Updated last year
- [ACL 2024] Code and data for "Machine Unlearning of Pre-trained Large Language Models"☆68Sep 30, 2024Updated last year
- [NeurIPS 2022] Explaining Graph Neural Networks with Structure-Aware Cooperative Games (GStarX)☆15Oct 20, 2022Updated 3 years ago
- Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging. Arxiv, 2024.☆16Oct 28, 2024Updated last year
- ☆45Mar 3, 2023Updated 3 years ago
- Butler 是一个用于自动化服务管理和任务调度的工具项目。☆17Updated this week