[EMNLP 2026] PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
☆18Aug 21, 2026Updated last week
Alternatives and similar repositories for PISanitizer
Users that are interested in PISanitizer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses☆27Updated this week
- [ACL 2026] PIArena: A Platform for Prompt Injection Evaluation☆46Apr 28, 2026Updated 4 months ago
- The official implementation of the paper "AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?"☆76May 19, 2026Updated 3 months ago
- Code for our NAACL2025 accepted paper: Attention Tracker: Detecting Prompt Injection Attacks in LLMs☆29Sep 19, 2025Updated 11 months ago
- Repo for the paper "Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks".☆70Jun 11, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15Jul 2, 2026Updated last month
- ASIDE: Architectural Separation of Instructions and Data in Language Models [ICLR 2026]☆17Jun 10, 2026Updated 2 months ago
- code of paper "Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM"☆14Nov 17, 2023Updated 2 years ago
- ☆26Jan 30, 2026Updated 7 months ago
- [NeurIPS 2025] The official implementation of the paper "DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agen…☆60Jul 16, 2026Updated last month
- ☆18Jun 19, 2023Updated 3 years ago
- Official release of code for the paper RL is a hammer and LLMs are nails A simple RL approach to stronger prompt injection attacks☆53May 6, 2026Updated 3 months ago
- Official implementation of the WASP web agent security benchmark☆98Apr 13, 2026Updated 4 months ago
- [S&P 2026] SoK: Evaluating Jailbreak Guardrails for Large Language Models☆46Dec 17, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆45May 29, 2026Updated 3 months ago
- [NDSS 2026] Official repo for Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography☆61Mar 14, 2026Updated 5 months ago
- [USENIX Security 2025] PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models☆298Jan 27, 2026Updated 7 months ago
- [ICLR'26 Oral] RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments☆60Feb 9, 2026Updated 6 months ago
- A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.☆777Jun 2, 2026Updated 2 months ago
- Official codes of KDD'24 paper "HiFGL: A Hierarchical Framework for Cross-silo Cross-device Federated Graph Learning"☆10Sep 4, 2024Updated last year
- ☆15Nov 29, 2025Updated 9 months ago
- DICE: Detecting In-distribution Data Contamination with LLM's Internal State☆11Sep 21, 2024Updated last year
- ☆12Sep 8, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- SkillGrad: Optimizing Agent Skills Like Gradient Descent☆26May 28, 2026Updated 3 months ago
- To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models☆34May 21, 2025Updated last year
- ICCV'2023: Combating Noisy Labels with Sample Selection by Mining High-Discrepancy Examples☆12Oct 16, 2023Updated 2 years ago
- Code&Data for the paper "Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents" [NeurIPS 2024]☆118Sep 27, 2024Updated last year
- Attribute statements generated by LLMs to preceding tokens using attention weights.☆28Apr 22, 2025Updated last year
- official implementation of [USENIX Sec'25] StruQ: Defending Against Prompt Injection with Structured Queries☆78Nov 10, 2025Updated 9 months ago
- ☆22Dec 11, 2025Updated 8 months ago
- [USENIX Security 2025] SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks☆23Sep 18, 2025Updated 11 months ago
- Proves that larger context windows don't fix RAG on structured data — they make wrong answers harder to detect. Then solves it with a que…☆15Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official implementation for "Mitigating Overthinking in Large Reasoning Models via Manifold Steering"☆15May 29, 2025Updated last year
- (ICLR 2026)Benchmarking LLM Tool-Use in the Wild☆39Apr 5, 2026Updated 4 months ago
- Code repository for the ICML 2026 Oral paper "Characterizing, Evaluating, and Optimizing Complex Reasoning".☆19Jun 21, 2026Updated 2 months ago
- [ACL 2025] LongSafety: Evaluating Long-Context Safety of Large Language Models☆16Jun 18, 2025Updated last year
- ☆40May 17, 2025Updated last year
- Code for Voice Jailbreak Attacks Against GPT-4o.☆39May 31, 2024Updated 2 years ago
- Official implementation for KDD'25 paper "Unleashing The Power of Pre-Trained Language Models for Irregularly Sampled Time Series".☆18Nov 28, 2025Updated 9 months ago