JailBench:大型语言模型越狱攻击风险评测中文数据集 [PAKDD 2025]
☆191Mar 3, 2025Updated last year
Alternatives and similar repositories for JailBench
Users that are interested in JailBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 🚀 JailbreakBench 是一个用于评估大语言模型(LLM)安全性的测试工具,专注于检测模型对越狱攻击(Jailbreak)的抵抗能力。通过模拟恶意提示词注入、编码攻击和多轮对话操控,量化模型的漏洞风险,并生成详细报告与可视化分析。支持中英文数据集,适用于安全研究…☆36Sep 1, 2025Updated 11 months ago
- ☆12Sep 29, 2024Updated last year
- "他山之 石、可以攻玉":复旦JADE团队发布的大模型测评与治理系列☆524Jul 23, 2026Updated 3 weeks ago
- SC-Safety: 中文大模型多轮对抗安全基准☆152Mar 15, 2024Updated 2 years ago
- 针对大语言模型的对抗性攻击总结☆38Dec 22, 2023Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- 针对大模型的后门攻击☆13Jun 30, 2024Updated 2 years ago
- Accepted by ECCV 2024☆218Oct 15, 2024Updated last year
- Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。☆1,195Feb 27, 2024Updated 2 years ago
- The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".☆18Jun 24, 2026Updated last month
- Flames is a highly adversarial benchmark in Chinese for LLM's harmlessness evaluation developed by Shanghai AI Lab and Fudan NLP Group.☆68May 21, 2024Updated 2 years ago
- Repo for the paper "Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks".☆70Jun 11, 2026Updated 2 months ago
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated 11 months ago
- [NeurIPS 2024] Fight Back Against Jailbreaking via Prompt Adversarial Tuning☆11Oct 29, 2024Updated last year
- ☆40May 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆30Updated this week
- ☆172Sep 2, 2024Updated last year
- ☆23Jul 26, 2025Updated last year
- 【ACL 2024】 SALAD benchmark & MD-Judge☆176Mar 8, 2025Updated last year
- Official implementation of “Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models” (AAAI 2026).☆37Mar 22, 2026Updated 4 months ago
- A ToB solution for flexibly detecting prompt injection risks across diverse open/closed AI infrastructure and packaged APIs.☆19Jun 26, 2025Updated last year
- ☆27Mar 17, 2025Updated last year
- The official repository for guided jailbreak benchmark☆32Jul 28, 2025Updated last year
- ☆134Jul 29, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A novel approach to improve the safety of large language models, enabling them to transition effectively from unsafe to safe state.☆72May 22, 2025Updated last year
- Official Repository for ACL 2024 Paper SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding☆153Jul 19, 2024Updated 2 years ago
- [ACL24] Official Repo of Paper `ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs`☆102Aug 15, 2025Updated last year
- A benchmark and generation framework for malicious agent skills.☆47Jun 10, 2026Updated 2 months ago
- SecProbe:任务驱动式大模型安全能力评测系统☆15Nov 29, 2024Updated last year
- It is a pure front-end tool for testing the security boundaries of large language models, helping researchers to find and fix potential s…☆21May 6, 2025Updated last year
- A reading list for large models safety, security, and privacy (including Awesome LLM Security, Safety, etc.).☆2,054Jun 17, 2026Updated last month
- LiveSecBench:动态中文大模型安全榜单☆29Mar 9, 2026Updated 5 months ago
- An easy-to-use Python framework to generate adversarial jailbreak prompts.☆884Mar 30, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆100Mar 20, 2025Updated last year
- Official Implementation of implicit reference attack☆11Oct 16, 2024Updated last year
- ☆67May 21, 2025Updated last year
- [ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language M…☆456Jan 22, 2025Updated last year
- Submission Guide + Discussion Board for AI Singapore Global Challenge for Safe and Secure LLMs (Track 1A).☆16Jul 4, 2024Updated 2 years ago
- [COLM 2024] JailBreakV-28K: A comprehensive benchmark designed to evaluate the transferability of LLM jailbreak attacks to MLLMs, and fur…☆96May 9, 2025Updated last year
- AmpleGCG: Learning a Universal and Transferable Generator of Adversarial Attacks on Both Open and Closed LLM☆87Nov 3, 2024Updated last year