JailBench:大型语言模型越狱攻击风险评测中文数据集 [PAKDD 2025]
☆191Mar 3, 2025Updated last year
Alternatives and similar repositories for JailBench
Users that are interested in JailBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 🚀 JailbreakBench 是一个用于评估大语言模型(LLM)安全性的测试工具,专注于检测模型对越狱攻击(Jailbreak)的抵抗能力。通过模拟恶意提示词注入、编码攻击和多轮对话操控,量化模型的漏洞风险,并生成详细报告与可视化分析。支持中英文数据集,适用于安全研究…☆37Sep 1, 2025Updated last year
- ☆12Sep 29, 2024Updated last year
- "他山之石、可以攻玉":复旦JADE团队发布的大模型 测评与治理系列☆527Jul 23, 2026Updated last month
- SC-Safety: 中文大模型多轮对抗安全基准☆152Mar 15, 2024Updated 2 years ago
- ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors [EMNLP 2024 Findings]☆230Sep 29, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 针对大语言模型的对抗性攻击总结☆38Dec 22, 2023Updated 2 years ago
- ☆18Jul 3, 2025Updated last year
- Accepted by ECCV 2024☆219Oct 15, 2024Updated last year
- Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。☆1,216Feb 27, 2024Updated 2 years ago
- The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".☆19Jun 24, 2026Updated 2 months ago
- Flames is a highly adversarial benchmark in Chinese for LLM's harmlessness evaluation developed by Shanghai AI Lab and Fudan NLP Group.☆68May 21, 2024Updated 2 years ago
- Repo for the paper "Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks".☆70Jun 11, 2026Updated 2 months ago
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated last year
- [NeurIPS 2024] Fight Back Against Jailbreaking via Prompt Adversarial Tuning☆11Oct 29, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆40May 17, 2025Updated last year
- ☆30Aug 12, 2026Updated 3 weeks ago
- ☆172Sep 2, 2024Updated 2 years ago
- ☆23Jul 26, 2025Updated last year
- 【ACL 2024】 SALAD benchmark & MD-Judge☆175Mar 8, 2025Updated last year
- [NDSS'25 Best Technical Poster] A collection of automated evaluators for assessing jailbreak attempts.☆196Apr 1, 2025Updated last year
- Official implementation of “Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models” (AAAI 2026).☆38Mar 22, 2026Updated 5 months ago
- A ToB solution for flexibly detecting prompt injection risks across diverse open/closed AI infrastructure and packaged APIs.☆19Jun 26, 2025Updated last year
- ☆27Mar 17, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- The official repository for guided jailbreak benchmark☆32Updated this week
- ☆132Jul 29, 2026Updated last month
- A novel approach to improve the safety of large language models, enabling them to transition effectively from unsafe to safe state.☆72May 22, 2025Updated last year
- Official Repository for ACL 2024 Paper SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding☆153Jul 19, 2024Updated 2 years ago
- Automatic Jailbreaking of the Text-to-Image Generative AI Systems☆15Jun 23, 2024Updated 2 years ago
- A benchmark and generation framework for malicious agent skills.☆57Jun 10, 2026Updated 2 months ago
- ☆15Apr 1, 2026Updated 5 months ago
- SecProbe:任务驱动式大模型安全能力评测系统☆15Nov 29, 2024Updated last year
- It is a pure front-end tool for testing the security boundaries of large language models, helping researchers to find and fix potential s…☆21May 6, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A reading list for large models safety, security, and privacy (including Awesome LLM Security, Safety, etc.).☆2,069Aug 28, 2026Updated last week
- An easy-to-use Python framework to generate adversarial jailbreak prompts.☆908Updated this week
- ☆102Mar 20, 2025Updated last year
- Official Implementation of implicit reference attack☆11Oct 16, 2024Updated last year
- ☆67May 21, 2025Updated last year
- [ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language M…☆461Jan 22, 2025Updated last year
- Submission Guide + Discussion Board for AI Singapore Global Challenge for Safe and Secure LLMs (Track 1A).☆16Jul 4, 2024Updated 2 years ago