agencyenterprise/PromptInject

Readme badge preview -

If you own this repo, copy the snippet below and add it to your README.md

[![RelatedRepos](https://img.shields.io/badge/related-repos-yellow)](https://relatedrepos.com/gh/agencyenterprise/PromptInject)

agencyenterprise / PromptInject

PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022

☆510

Alternatives and similar repositories for PromptInject

Users that are interested in PromptInject are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

sherdencooper / GPTFuzz
View on GitHub
Official repo for GPTFUZZER : Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
☆601Feb 27, 2026Updated 4 months ago
greshake / llm-security
View on GitHub
New ways of breaking app-integrated LLMs
☆2,118Jul 17, 2025Updated last year
llm-attacks / llm-attacks
View on GitHub
Universal and Transferable Attacks on Aligned Language Models
☆4,740Aug 2, 2024Updated last year
corca-ai / awesome-llm-security
View on GitHub
A curation of awesome tools, documents and projects about LLM Security.
☆1,659Aug 20, 2025Updated 11 months ago
SheltonLiu-N / AutoDAN
View on GitHub
[ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language M…
☆453Jan 22, 2025Updated last year
GPUs on demand by Runpod - Special Offer Available • Ad
Run AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
patrickrchao / JailbreakingLLMs
View on GitHub
☆756Jul 2, 2025Updated last year
mnns / LLMFuzzer
View on GitHub
🧠 LLMFuzzer - Fuzzing Framework for Large Language Models 🧠 LLMFuzzer is the first open-source fuzzing framework specifically designed …
☆364Feb 12, 2024Updated 2 years ago
LLM-Canary / LLM-Canary
View on GitHub
☆29Feb 5, 2024Updated 2 years ago
NVIDIA / garak
View on GitHub
the LLM vulnerability scanner
☆8,510Jul 14, 2026Updated last week
hwchase17 / adversarial-prompts
View on GitHub
Curation of prompts that are known to be adversarial to large language models
☆193Feb 12, 2023Updated 3 years ago
liu00222 / Open-Prompt-Injection
View on GitHub
This repository provides a benchmark for prompt injection attacks and defenses in LLMs
☆465Oct 29, 2025Updated 8 months ago
protectai / rebuff
View on GitHub
LLM Prompt Injection Detector
☆1,513Aug 7, 2024Updated last year
centerforaisafety / HarmBench
View on GitHub
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
☆1,011Aug 16, 2024Updated last year
LLMSecurity / HouYi
View on GitHub
The automated prompt injection framework for LLM-integrated applications.
☆269Sep 12, 2024Updated last year
Managed hosting for WordPress and PHP on Cloudways • Ad
Managed hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
deadbits / vigil-llm
View on GitHub
⚡ Vigil ⚡ Detect prompt injections, jailbreaks, and other potentially risky Large Language Model (LLM) inputs
☆491Jan 31, 2024Updated 2 years ago
Algorithmic-Alignment-Lab / CommonClaim
View on GitHub
Explore, Establish, Exploit: Red Teaming Language Models from Scratch
☆15Jun 21, 2023Updated 3 years ago
utkusen / promptmap
View on GitHub
a security scanner for custom LLM applications
☆1,230Dec 1, 2025Updated 7 months ago
declare-lab / red-instruct
View on GitHub
Codes and datasets of the paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
☆111Mar 8, 2024Updated 2 years ago
egozverev / aside
View on GitHub
ASIDE: Architectural Separation of Instructions and Data in Language Models [ICLR 2026]
☆16Jun 10, 2026Updated last month
sleeepeer / PISanitizer
View on GitHub
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
☆18Dec 10, 2025Updated 7 months ago
tml-epfl / llm-adaptive-attacks
View on GitHub
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks [ICLR 2025]
☆391Jan 23, 2025Updated last year
ejones313 / auditing-llms
View on GitHub
☆61Mar 9, 2023Updated 3 years ago
dropbox / llm-security
View on GitHub
Dropbox LLM Security research code and results
☆259May 21, 2024Updated 2 years ago
Deploy on Railway without the complexity - Free Credits Offer • Ad
Connect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
LLM-Tuning-Safety / LLMs-Finetuning-Safety
View on GitHub
We jailbreak GPT-3.5 Turbo’s safety guardrails by fine-tuning it on only 10 adversarially designed examples, at a cost of less than $0.20…
☆358Feb 23, 2024Updated 2 years ago
vinusankars / BEAST
View on GitHub
Implementation of BEAST adversarial attack for language models (ICML 2024)
☆89May 14, 2024Updated 2 years ago
alphadl / SafeLLM_with_IntentionAnalysis
View on GitHub
Towards Safe LLM with our simple-yet-highly-effective Intention Analysis Prompting
☆21Mar 25, 2024Updated 2 years ago
papersPapers / BadPrompt
View on GitHub
Code for the paper "BadPrompt: Backdoor Attacks on Continuous Prompts"
☆41Jul 8, 2024Updated 2 years ago
JailbreakBench / jailbreakbench
View on GitHub
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models [NeurIPS 2024 Datasets and Benchmarks Track]
☆634Apr 4, 2025Updated last year
Princeton-SysML / Jailbreak_LLM
View on GitHub
☆203Nov 26, 2023Updated 2 years ago
facebookresearch / SecAlign
View on GitHub
Repo for the research paper "SecAlign: Defending Against Prompt Injection with Preference Optimization"
☆98Jul 2, 2026Updated 2 weeks ago
alexandrasouly / strongreject
View on GitHub
Repository for "StrongREJECT for Empty Jailbreaks" paper
☆157Nov 3, 2024Updated last year
chawins / llm-sp
View on GitHub
Papers and resources related to the security and privacy of LLMs 🤖
☆579Jun 8, 2025Updated last year
Deploy to Railway using AI coding agents - Free Credits Offer • Ad
Use Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
leondz / lm_risk_cards
View on GitHub
Risks and targets for assessing LLMs & LLM vulnerabilities
☆36May 27, 2024Updated 2 years ago
RICommunity / TAP
View on GitHub
TAP: An automated jailbreaking method for black-box LLMs
☆241Dec 10, 2024Updated last year
ebagdasa / multimodal_injection
View on GitHub
☆101Oct 15, 2023Updated 2 years ago
tldrsec / prompt-injection-defenses
View on GitHub
Every practical and proposed defense against prompt injection.
☆713Feb 22, 2025Updated last year
anthropics / hh-rlhf
View on GitHub
Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"
☆1,851Jun 17, 2025Updated last year
leondz / autoredteam
View on GitHub
autoredteam: code for training models that automatically red team other language models
☆17Aug 9, 2023Updated 2 years ago
RobustNLP / CipherChat
View on GitHub
A framework to evaluate the generalization capability of safety alignment for LLMs
☆628Oct 9, 2025Updated 9 months ago