JailBench:大型语言模型越狱攻击风险评测中文数据集 [PAKDD 2025]
☆193Mar 3, 2025Updated last year
Alternatives and similar repositories for JailBench
Users that are interested in JailBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 🚀 JailbreakBench 是一个用于评估大语言模型(LLM)安全性的测试工具,专注于检测模型对越狱攻击(Jailbreak)的抵抗能力。通过模拟恶意提示词注入、编码攻击和多轮对话操控,量化模型的漏洞风险,并生成详细报告与可视化分析。支持中英文数据集,适用于安全研究…☆38Sep 1, 2025Updated last year
- ☆12Sep 29, 2024Updated last year
- "他山之石、可以攻玉": 复旦JADE团队发布的大模型测评与治理系列☆528Jul 23, 2026Updated 2 months ago
- ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors [EMNLP 2024 Findings]☆231Sep 29, 2024Updated last year
- 针对大语言模型的对抗性攻击总结☆38Dec 22, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆18Jul 3, 2025Updated last year
- 针对大模型的后门攻击☆13Jun 30, 2024Updated 2 years ago
- Accepted by ECCV 2024☆221Oct 15, 2024Updated last year
- Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。☆1,219Feb 27, 2024Updated 2 years ago
- The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".☆19Jun 24, 2026Updated 3 months ago
- Flames is a highly adversarial benchmark in Chinese for LLM's harmlessness evaluation developed by Shanghai AI Lab and Fudan NLP Group.☆68May 21, 2024Updated 2 years ago
- Repo for the paper "Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks".☆71Jun 11, 2026Updated 3 months ago
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated last year
- [NeurIPS 2024] Fight Back Against Jailbreaking via Prompt Adversarial Tuning☆11Oct 29, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆40May 17, 2025Updated last year
- ☆30Aug 12, 2026Updated last month
- ☆172Sep 2, 2024Updated 2 years ago
- ☆23Jul 26, 2025Updated last year
- 【ACL 2024】 SALAD benchmark & MD-Judge☆175Mar 8, 2025Updated last year
- [NDSS'25 Best Technical Poster] A collection of automated evaluators for assessing jailbreak attempts.☆198Apr 1, 2025Updated last year
- Official implementation of “Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models” (AAAI 2026).☆39Mar 22, 2026Updated 6 months ago
- A ToB solution for flexibly detecting prompt injection risks across diverse open/closed AI infrastructure and packaged APIs.☆19Jun 26, 2025Updated last year
- ☆27Mar 17, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The official repository for guided jailbreak benchmark☆32Aug 31, 2026Updated 3 weeks ago
- ☆134Jul 29, 2026Updated last month
- A novel approach to improve the safety of large language models, enabling them to transition effectively from unsafe to safe state.☆72May 22, 2025Updated last year
- Official Repository for ACL 2024 Paper SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding☆153Jul 19, 2024Updated 2 years ago
- [ACL24] Official Repo of Paper `ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs`☆100Aug 15, 2025Updated last year
- Automatic Jailbreaking of the Text-to-Image Generative AI Systems☆15Jun 23, 2024Updated 2 years ago
- A benchmark and generation framework for malicious agent skills.☆60Jun 10, 2026Updated 3 months ago
- ☆15Apr 1, 2026Updated 5 months ago
- It is a pure front-end tool for testing the security boundaries of large language models, helping researchers to find and fix potential s…☆21May 6, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An easy-to-use Python framework to generate adversarial jailbreak prompts.☆911Sep 1, 2026Updated 3 weeks ago
- ☆103Mar 20, 2025Updated last year
- Official Implementation of implicit reference attack☆11Oct 16, 2024Updated last year
- ☆68May 21, 2025Updated last year
- [ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language M…☆465Jan 22, 2025Updated last year
- Submission Guide + Discussion Board for AI Singapore Global Challenge for Safe and Secure LLMs (Track 1A).☆16Jul 4, 2024Updated 2 years ago
- [COLM 2024] JailBreakV-28K: A comprehensive benchmark designed to evaluate the transferability of LLM jailbreak attacks to MLLMs, and fur…☆99May 9, 2025Updated last year