An easy-to-use Python framework to defend against jailbreak prompts.
☆21Mar 22, 2025Updated last year
Alternatives and similar repositories for llm-defense
Users that are interested in llm-defense are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A novel jailbreak attack unveiling an overlooked attack surface inherently in the chain-of-thought reasoning trajectory of LLMs☆22Apr 3, 2026Updated 4 months ago
- Towards Safe LLM with our simple-yet-highly-effective Intention Analysis Prompting☆21Mar 25, 2024Updated 2 years ago
- A curated collection of research and techniques for protecting intellectual property of large language models, including watermarking, fi…☆52Jun 10, 2026Updated 2 months ago
- Red Queen Dataset and data generation template☆27Dec 26, 2025Updated 7 months ago
- PyTorch Implementation of the paper "Defining and Quantifying the Emergence of Sparse Concepts in DNNs" (CVPR 2023)☆12Dec 24, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆11May 7, 2022Updated 4 years ago
- ☆12Mar 11, 2025Updated last year
- SVIP: Towards Verifiable Inference of Open-Source Large Language Models☆15Jun 3, 2025Updated last year
- ☆12Sep 29, 2024Updated last year
- Understanding the Robustness of Skeleton-based Action Recognition under Adversarial Attack CVPR 2021☆15Mar 8, 2024Updated 2 years ago
- BASAR:Black-box Attack on Skeletal Action Recognition, CVPR 2021☆19Feb 18, 2025Updated last year
- Understanding Why and How Instruction Tuning Changes Pre-trained Models☆25Mar 18, 2024Updated 2 years ago
- Implementation of paper 'Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing'☆24Jun 9, 2024Updated 2 years ago
- 北京航空航天大学课程资料共享仓库☆10Apr 21, 2019Updated 7 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementation for "RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content"☆24Jul 28, 2024Updated 2 years ago
- Everything about xss protection technology☆15Oct 22, 2019Updated 6 years ago
- ☆15Jun 2, 2026Updated 2 months ago
- ☆13Mar 29, 2021Updated 5 years ago
- ☆13Nov 11, 2022Updated 3 years ago
- ICL backdoor attack☆17Nov 4, 2024Updated last year
- Some vulnerability research slides that I made☆12Jan 5, 2022Updated 4 years ago
- my poc☆16Oct 28, 2020Updated 5 years ago
- URL-encode data streams via commandline☆14Oct 26, 2019Updated 6 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Mobile Security - OMTG-Android Walkthrough☆11Oct 31, 2019Updated 6 years ago
- Unofficial implementation of "Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection"☆27Jul 6, 2024Updated 2 years ago
- ☆35Nov 12, 2024Updated last year
- Tool to get the top android apps for bug bounty purpose☆17Sep 10, 2020Updated 5 years ago
- [ECCV2024] Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajector…☆32Nov 15, 2025Updated 8 months ago
- Agent3σ-Canary is an evaluation framework for AI Agent security in realistic runtime environments.☆34Updated this week
- Can Knowledge Editing Really Correct Hallucinations? (ICLR 2025)☆27Aug 10, 2025Updated last year
- PitchVC: Pitch Conditioned Any-to-Many Voice Conversion☆35Jun 6, 2024Updated 2 years ago
- Own collection dictionary☆14Apr 20, 2020Updated 6 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ESEC/FSE'21: Prediction-Preserving Program Simplification☆10Oct 4, 2022Updated 3 years ago
- LLM Self Defense: By Self Examination, LLMs know they are being tricked☆52May 21, 2024Updated 2 years ago
- Code and dataset for the paper: "Can Editing LLMs Inject Harm?" [AAAI'26]☆21Dec 26, 2025Updated 7 months ago
- some codeql rules☆15Apr 6, 2020Updated 6 years ago
- A substitute repository put up on public demand for the original Awesome WAF repository (https://github.com/0xInfection/Awesome-WAF) whic…☆14May 3, 2019Updated 7 years ago
- 毕业设计。Keywords: 层次聚类、谱聚类、WordNet☆10Jun 29, 2014Updated 12 years ago
- ASR (Automatic Speech Recognition) for real-time streamed audio powered by Whisper and tranformers☆36Apr 22, 2026Updated 3 months ago