Papers about red teaming LLMs and Multimodal models.
☆179Jul 14, 2026Updated 2 months ago
Alternatives and similar repositories for OpenRedTeaming
Users that are interested in OpenRedTeaming are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆23May 20, 2025Updated last year
- A curated list of awesome LLM Red Teaming training, resources, and tools.☆134Sep 4, 2025Updated last year
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal☆1,052Aug 16, 2024Updated 2 years ago
- Simple Chatbot for testing AI Red Team tooling☆16Feb 11, 2025Updated last year
- [ACL 25] SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities☆30Apr 2, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of ICLR'24 paper, "Curiosity-driven Red Teaming for Large Language Models" (https://openreview.net/pdf?id=4KqkizX…☆91Mar 15, 2024Updated 2 years ago
- Security Threats related with MCP (Model Context Protocol), MCP Servers and more☆51Apr 24, 2025Updated last year
- An automated pipeline that leverages LLM's meta-learning capability to iteratively design and refine red-teaming systems without human in…☆31May 24, 2026Updated 3 months ago
- Do you want to learn AI Security but don't know where to start ? Take a look at this map.☆29Apr 23, 2024Updated 2 years ago
- An easy-to-use Python framework to generate adversarial jailbreak prompts.☆911Sep 1, 2026Updated 3 weeks ago
- An ongoing & curated collection of awesome vulnerability scanning software, libraries and frameworks, best guidelines, technical resource…☆14Feb 7, 2022Updated 4 years ago
- [NAACL 2025] The official implementation of paper "Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language M…☆28Mar 14, 2024Updated 2 years ago
- ☆32Feb 23, 2025Updated last year
- Papers and resources related to the security and privacy of LLMs 🤖☆583Jun 8, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language M…☆465Jan 22, 2025Updated last year
- A reading list for large models safety, security, and privacy (including Awesome LLM Security, Safety, etc.).☆2,080Sep 2, 2026Updated 2 weeks ago
- We jailbreak GPT-3.5 Turbo’s safety guardrails by fine-tuning it on only 10 adversarially designed examples, at a cost of less than $0.20…☆359Feb 23, 2024Updated 2 years ago
- [ICML 2025] UDora: A Unified Red Teaming Framework against LLM Agents☆39Jun 24, 2025Updated last year
- [NDSS'25 Best Technical Poster] A collection of automated evaluators for assessing jailbreak attempts.☆198Apr 1, 2025Updated last year
- [EMNLP 2024] Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction☆17Nov 9, 2024Updated last year
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks [ICLR 2025]☆395Jan 23, 2025Updated last year
- FeatureAlignment = Alignment + Mechanistic Interpretability☆35Mar 8, 2025Updated last year
- An implementation for MLLM oversensitivity evaluation☆18Nov 16, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆790Jul 2, 2025Updated last year
- LobotoMl is a set of scripts and tools to assess production deployments of ML services☆10May 16, 2022Updated 4 years ago
- Our research proposes a novel MoGU framework that improves LLMs' safety while preserving their usability.☆18Jan 14, 2025Updated last year
- ☆24Jun 13, 2024Updated 2 years ago
- AIBOM Workshop RSA 2024☆15May 20, 2024Updated 2 years ago
- This repository provides a benchmark for prompt injection attacks and defenses in LLMs☆499Sep 12, 2026Updated last week
- The respository describing a novel datasets for word association explanations☆13Sep 21, 2023Updated 3 years ago
- A curation of awesome tools, documents and projects about LLM Security.☆1,702Aug 20, 2025Updated last year
- ☆56Sep 9, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Röttger et al. (NAACL 2024): "XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models"☆147Feb 24, 2025Updated last year
- TACL 2025: Investigating Adversarial Trigger Transfer in Large Language Models☆20Aug 17, 2025Updated last year
- Official Repository for ACL 2024 Paper SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding☆153Jul 19, 2024Updated 2 years ago
- Benchmark evaluation code for "SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal" (ICLR 2025)☆87Mar 1, 2025Updated last year
- ☆17Jun 11, 2026Updated 3 months ago
- A curated list of awesome resources about LLM supply chain security (including papers, security reports and CVEs)☆107Jan 20, 2025Updated last year
- ☆24Dec 18, 2025Updated 9 months ago