The official repository for guided jailbreak benchmark
☆30Jul 28, 2025Updated 11 months ago
Alternatives and similar repositories for AI-Safety_Benchmark
Users that are interested in AI-Safety_Benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code Implementation of Adversarial Prompt Evaluation paper☆14Sep 18, 2025Updated 9 months ago
- ☆22Oct 25, 2024Updated last year
- A simple academic poster template using pptx, for NeurIPS/ICML/ICCV.☆46Aug 1, 2025Updated 11 months ago
- ☆136Updated this week
- A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack–Defense Evaluation☆74May 8, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official implementation of paper: DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers☆69Aug 25, 2024Updated last year
- A fast + lightweight implementation of the GCG algorithm in PyTorch☆340May 13, 2025Updated last year
- [ICML 2025] An official source code for paper "FlipAttack: Jailbreak LLMs via Flipping".☆175May 2, 2025Updated last year
- [USENIX'25] HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns☆14Mar 1, 2025Updated last year
- [ECCV'24 Oral] The official GitHub page for ''Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking …☆38Oct 23, 2024Updated last year
- Improved techniques for optimization-based jailbreaking on large language models (ICLR2025)☆146Apr 7, 2025Updated last year
- Test LLMs against jailbreaks and unprecedented harms☆41Oct 19, 2024Updated last year
- [USENIX Security'24] Official repository of "Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise a…☆115Oct 11, 2024Updated last year
- Accept by CVPR 2025 (highlight)☆25Jun 8, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆25Jan 17, 2025Updated last year
- This is the code repository for "Uncovering Safety Risks of Large Language Models through Concept Activation Vector"☆49Oct 13, 2025Updated 8 months ago
- Code for the paper "Jailbreak Large Vision-Language Models Through Multi-Modal Linkage"☆35Dec 6, 2024Updated last year
- [ICLR 2026] The official code for "Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models"☆28Feb 7, 2026Updated 4 months ago
- ☆22May 14, 2025Updated last year
- ☆67May 21, 2025Updated last year
- Panda Guard is designed for researching jailbreak attacks, defenses, and evaluation algorithms for large language models (LLMs).☆68Mar 23, 2026Updated 3 months ago
- Code repo of our paper Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis (https://arxiv.org/abs/2406.10794…☆24Jul 26, 2024Updated last year
- Revisiting Character-level Adversarial Attacks for Language Models, ICML 2024☆19Feb 12, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- official implementation of [USENIX Sec'25] StruQ: Defending Against Prompt Injection with Structured Queries☆76Nov 10, 2025Updated 7 months ago
- ☆33Jun 24, 2024Updated 2 years ago
- [ICLR 2025] BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks☆31Nov 2, 2025Updated 8 months ago
- Do you want to learn AI Security but don't know where to start ? Take a look at this map.☆31Apr 23, 2024Updated 2 years ago
- Revolve: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization☆22Dec 13, 2024Updated last year
- DSN jailbreak Attack & Evaluation Ensemble☆17Feb 7, 2026Updated 4 months ago
- This repo is the artifact of FUEL☆16May 19, 2026Updated last month
- Code repository for the paper "Heuristic Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models"☆19Aug 7, 2025Updated 10 months ago
- Code for "When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search" (NeurIPS 2024)☆18Oct 22, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆138Dec 3, 2025Updated 7 months ago
- [NDSS'25 Best Technical Poster] A collection of automated evaluators for assessing jailbreak attempts.☆195Apr 1, 2025Updated last year
- ☆60Jun 5, 2024Updated 2 years ago
- [NeurIPS 2024] Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling☆35Nov 8, 2024Updated last year
- ☆70Jun 1, 2025Updated last year
- [NeurIPS 2021] "Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models" by Boxin Wang*, Chejian Xu*, Shuoh…☆13Apr 3, 2023Updated 3 years ago
- 网络安全 LLM 智能体应用教程☆30Mar 2, 2025Updated last year