A curated list of safety-related papers, articles, and resources focused on Large Language Models (LLMs). This repository aims to provide researchers, practitioners, and enthusiasts with insights into the safety implications, challenges, and advancements surrounding these powerful models.
☆1,904Jul 12, 2026Updated last month
Alternatives and similar repositories for Awesome-LLM-Safety
Users that are interested in Awesome-LLM-Safety are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A reading list for large models safety, security, and privacy (including Awesome LLM Security, Safety, etc.).☆2,069Aug 28, 2026Updated last week
- A curation of awesome tools, documents and projects about LLM Security.☆1,690Aug 20, 2025Updated last year
- ☆60Jun 13, 2024Updated 2 years ago
- An easy-to-use Python framework to generate adversarial jailbreak prompts.☆908Updated this week
- Papers and resources related to the security and privacy of LLMs 🤖☆582Jun 8, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [NAACL2024] Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey☆112Aug 7, 2024Updated 2 years ago
- Universal and Transferable Attacks on Aligned Language Models☆4,780Aug 2, 2024Updated 2 years ago
- 😎 up-to-date & curated list of awesome Attacks on Large-Vision-Language-Models papers, methods & resources.☆576Updated this week
- Accepted by IJCAI-24 Survey Track☆234Aug 25, 2024Updated 2 years ago
- A survey on harmful fine-tuning attack for large language model (ACM CSUR)☆246Jun 22, 2026Updated 2 months ago
- We jailbreak GPT-3.5 Turbo’s safety guardrails by fine-tuning it on only 10 adversarially designed examples, at a cost of less than $0.20…☆358Feb 23, 2024Updated 2 years ago
- [ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language M…☆461Jan 22, 2025Updated last year
- Awesome-Jailbreak-on-LLMs is a collection of state-of-the-art, novel, exciting jailbreak methods on LLMs. It contains papers, codes, data…☆1,611Aug 10, 2026Updated 3 weeks ago
- ☆197Oct 31, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety☆286Apr 12, 2026Updated 4 months ago
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models [NeurIPS 2024 Datasets and Benchmarks Track]☆664Apr 4, 2025Updated last year
- Official Repository for ACL 2024 Paper SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding☆153Jul 19, 2024Updated 2 years ago
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal☆1,037Aug 16, 2024Updated 2 years ago
- A lightweight library for large laguage model (LLM) jailbreaking defense.☆61Sep 11, 2025Updated 11 months ago
- Awesome Large Reasoning Model(LRM) Safety.This repository is used to collect security-related research on large reasoning models such as …☆84Updated this week
- ☆776Jul 2, 2025Updated last year
- ☆81Jan 21, 2026Updated 7 months ago
- A fast + lightweight implementation of the GCG algorithm in PyTorch☆353May 13, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Repository for the Paper (AAAI 2024, Oral) --- Visual Adversarial Examples Jailbreak Large Language Models☆282May 13, 2024Updated 2 years ago
- Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。☆1,216Feb 27, 2024Updated 2 years ago
- Official Repository for The Paper: Safety Alignment Should Be Made More Than Just a Few Tokens Deep☆191Apr 23, 2025Updated last year
- Official repository for "Safety in Large Reasoning Models: A Survey" - Exploring safety risks, attacks, and defenses for Large Reasoning …☆90Aug 25, 2025Updated last year
- [AAAI'25 (Oral)] Jailbreaking Large Vision-language Models via Typographic Visual Prompts☆213Jun 26, 2025Updated last year
- [ICML 2024] TrustLLM: Trustworthiness in Large Language Models☆632Jun 24, 2025Updated last year
- 【ACL 2024】 SALAD benchmark & MD-Judge☆175Mar 8, 2025Updated last year
- Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs. Empirical tricks for LLM Jailbreaking. (NeurIPS 2024)☆168Nov 30, 2024Updated last year
- A Survey on Jailbreak Attacks and Defenses against Multimodal Generative Models☆334Jan 11, 2026Updated 7 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official repo for GPTFUZZER : Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts☆608Feb 27, 2026Updated 6 months ago
- ☆204Nov 26, 2023Updated 2 years ago
- Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback☆1,613Nov 24, 2025Updated 9 months ago
- Official github repo for SafetyBench, a comprehensive benchmark to evaluate LLMs' safety. [ACL 2024]☆297Jul 28, 2025Updated last year
- Accepted by ECCV 2024☆219Oct 15, 2024Updated last year
- ☆172Sep 2, 2024Updated 2 years ago
- Awesome-LLM-Robustness: a curated list of Uncertainty, Reliability and Robustness in Large Language Models☆835Jun 5, 2026Updated 2 months ago