JailBench:大型语言模型越狱攻击风险评测中文数据集 [PAKDD 2025]
☆191Mar 3, 2025Updated last year
Alternatives and similar repositories for JailBench
Users that are interested in JailBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 🚀 JailbreakBench 是一个用于评估大语言模型(LLM)安全性的测试工具,专注于检测模型对越狱攻击(Jailbreak)的抵抗能力。通过模拟恶意提示词注入、编码攻击和多轮对话操控,量化模型的漏洞风险,并生成详细报告与可视化分析。支持中英文数据集,适用于安全研究…☆37Sep 1, 2025Updated 10 months ago
- ☆12Sep 29, 2024Updated last year
- "他山之石、可以攻玉":复旦JADE团队发布的大模型测评与治理系列☆520Updated this week
- SC-Safety: 中文大模型多轮对抗安全基准☆151Mar 15, 2024Updated 2 years ago
- ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors [EMNLP 2024 Findings]☆231Sep 29, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 针对大语言模型的对抗性攻击总结☆38Dec 22, 2023Updated 2 years ago
- ☆18Jul 3, 2025Updated last year
- Accepted by ECCV 2024☆218Oct 15, 2024Updated last year
- Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。☆1,189Feb 27, 2024Updated 2 years ago
- The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".☆18Jun 24, 2026Updated last month
- Flames is a highly adversarial benchmark in Chinese for LLM's harmlessness evaluation developed by Shanghai AI Lab and Fudan NLP Group.☆68May 21, 2024Updated 2 years ago
- Repo for the paper "Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks".☆70Jun 11, 2026Updated last month
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated 10 months ago
- [NeurIPS 2024] Fight Back Against Jailbreaking via Prompt Adversarial Tuning☆11Oct 29, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆40May 17, 2025Updated last year
- ☆30Mar 20, 2024Updated 2 years ago
- ☆172Sep 2, 2024Updated last year
- ☆22Jul 26, 2025Updated 11 months ago
- [NDSS'25 Best Technical Poster] A collection of automated evaluators for assessing jailbreak attempts.☆193Apr 1, 2025Updated last year
- Official implementation of “Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models” (AAAI 2026).☆37Mar 22, 2026Updated 4 months ago
- A ToB solution for flexibly detecting prompt injection risks across diverse open/closed AI infrastructure and packaged APIs.☆19Jun 26, 2025Updated last year
- ☆27Mar 17, 2025Updated last year
- The official repository for guided jailbreak benchmark☆31Jul 28, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆134Jun 29, 2026Updated 3 weeks ago
- A novel approach to improve the safety of large language models, enabling them to transition effectively from unsafe to safe state.☆72May 22, 2025Updated last year
- Official Repository for ACL 2024 Paper SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding☆154Jul 19, 2024Updated 2 years ago
- [ACL24] Official Repo of Paper `ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs`☆102Aug 15, 2025Updated 11 months ago
- Automatic Jailbreaking of the Text-to-Image Generative AI Systems☆15Jun 23, 2024Updated 2 years ago
- A benchmark and generation framework for malicious agent skills.☆40Jun 10, 2026Updated last month
- ☆15Apr 1, 2026Updated 3 months ago
- SecProbe:任务驱动式大模型安全能力评测系统☆15Nov 29, 2024Updated last year
- It is a pure front-end tool for testing the security boundaries of large language models, helping researchers to find and fix potential s…☆21May 6, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A reading list for large models safety, security, and privacy (including Awesome LLM Security, Safety, etc.).☆2,020Jun 17, 2026Updated last month
- LiveSecBench:动态中文大模型安全榜单☆29Mar 9, 2026Updated 4 months ago
- An easy-to-use Python framework to generate adversarial jailbreak prompts.☆874Mar 30, 2026Updated 3 months ago
- ☆99Mar 20, 2025Updated last year
- Official Implementation of implicit reference attack☆11Oct 16, 2024Updated last year
- ☆67May 21, 2025Updated last year
- [ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language M…☆453Jan 22, 2025Updated last year