An easy-to-use Python framework to defend against jailbreak prompts.
☆21Mar 22, 2025Updated last year
Alternatives and similar repositories for llm-defense
Users that are interested in llm-defense are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A novel jailbreak attack unveiling an overlooked attack surface inherently in the chain-of-thought reasoning trajectory of LLMs☆22Apr 3, 2026Updated 5 months ago
- A curated collection of research and techniques on securing agent communication in Large Language Model (LLM)–based agent systems, includ…☆28Dec 2, 2025Updated 9 months ago
- Red Queen Dataset and data generation template☆29Dec 26, 2025Updated 8 months ago
- Code for paper "Defending aginast LLM Jailbreaking via Backtranslation"☆34Aug 16, 2024Updated 2 years ago
- PyTorch Implementation of the paper "Defining and Quantifying the Emergence of Sparse Concepts in DNNs" (CVPR 2023)☆12Dec 24, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆11May 7, 2022Updated 4 years ago
- ☆12Mar 11, 2025Updated last year
- SVIP: Towards Verifiable Inference of Open-Source Large Language Models☆15Jun 3, 2025Updated last year
- ☆12Sep 29, 2024Updated last year
- Understanding the Robustness of Skeleton-based Action Recognition under Adversarial Attack CVPR 2021☆15Mar 8, 2024Updated 2 years ago
- [CVPR 2024] "Transferable Structural Sparse Adversarial Attack Via Exact Group Sparsity Training", Di Ming, Peng Ren, Yunlong Wang, Xin …☆16Jun 11, 2024Updated 2 years ago
- BASAR:Black-box Attack on Skeletal Action Recognition, CVPR 2021☆19Feb 18, 2025Updated last year
- Code for SLT 2016 paper on Grapheme-to-Phoneme conversion using attention based encoder-decoder models☆15Feb 20, 2019Updated 7 years ago
- [EMNLP 2023] Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-Thoughts☆28Nov 4, 2023Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Implementation of paper 'Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing'☆24Jun 9, 2024Updated 2 years ago
- 北京航空航天大学课程资料共享仓库☆10Apr 21, 2019Updated 7 years ago
- Implementation for "RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content"☆24Jul 28, 2024Updated 2 years ago
- ☆24Jun 16, 2024Updated 2 years ago
- DILMA: Differentiable Language Model Adversarial Attacks on Categorical Sequence Classifiers☆12Oct 7, 2020Updated 5 years ago
- Code for NDSS paper: Stealthy Adversarial Perturbations Against Real-Time Video Classification Systems☆21Nov 24, 2018Updated 7 years ago
- ☆15Jun 2, 2026Updated 3 months ago
- ☆13Mar 29, 2021Updated 5 years ago
- ☆13Nov 11, 2022Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This is the official code implementation of paper De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks.☆16Sep 10, 2025Updated last year
- A compact toolbox for backdoor attacks and defenses.☆194Jul 16, 2024Updated 2 years ago
- ICL backdoor attack☆17Nov 4, 2024Updated last year
- Some vulnerability research slides that I made☆12Jan 5, 2022Updated 4 years ago
- ☆27Feb 22, 2024Updated 2 years ago
- URL-encode data streams via commandline☆14Oct 26, 2019Updated 6 years ago
- ☆13Dec 22, 2017Updated 8 years ago
- Mobile Security - OMTG-Android Walkthrough☆11Oct 31, 2019Updated 6 years ago
- [CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment☆29Jun 11, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Unified Approach to Interpreting and Boosting Adversarial Transferability (ICLR2021)☆32Apr 22, 2022Updated 4 years ago
- Evaluation Metrics Used For The Performance Evaluation of Voice Conversion (VC) Models☆19Jul 8, 2025Updated last year
- Unofficial implementation of "Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection"☆27Jul 6, 2024Updated 2 years ago
- ☆35Nov 12, 2024Updated last year
- Robustify Black-Box Models (ICLR'22 - Spotlight)☆23Jan 29, 2023Updated 3 years ago
- Tool to get the top android apps for bug bounty purpose☆17Sep 10, 2020Updated 6 years ago
- [ECCV2024] Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajector…☆32Nov 15, 2025Updated 10 months ago