Code for our NeurIPS 2024 paper Improved Generation of Adversarial Examples Against Safety-aligned LLMs
☆12Nov 7, 2024Updated last year
Alternatives and similar repositories for Gradient-based-Jailbreak-Attacks
Users that are interested in Gradient-based-Jailbreak-Attacks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for Findings-EMNLP 2023 paper: Multi-step Jailbreaking Privacy Attacks on ChatGPT☆37Oct 15, 2023Updated 2 years ago
- A repo for LLM jailbreak☆14Sep 5, 2023Updated 2 years ago
- [Tensorflow] A Game Theoretic approach using GAN for Phishing URL synthesis and detection☆11Nov 14, 2022Updated 3 years ago
- All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks☆17Apr 24, 2024Updated 2 years ago
- [ACL 25] SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities☆30Apr 2, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This novel adversarial attack method that I have developed is called CEIA (Contextual Embedding Inversion Attack). This method is a sophi…☆16Nov 19, 2024Updated last year
- Repository for project codes related to turbulence modelling.☆17Sep 30, 2024Updated last year
- [CIKM 2024] Trojan Activation Attack: Attack Large Language Models using Activation Steering for Safety-Alignment.☆30Jul 29, 2024Updated last year
- ☆17Mar 8, 2024Updated 2 years ago
- ☆13Jan 14, 2025Updated last year
- 2021年暨南大学CTF新生赛题目与源码☆15Dec 6, 2021Updated 4 years ago
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- Identification of the Adversary from a Single Adversarial Example (ICML 2023)☆10Jul 15, 2024Updated 2 years ago
- Links to publications that focus on the interpretation and analysis of in-context learning☆14Oct 17, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICLR 2026] Rectifying LLM Thought From Lens of Optimization☆15Dec 5, 2025Updated 7 months ago
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks [ICLR 2025]☆391Jan 23, 2025Updated last year
- ☆10Apr 28, 2020Updated 6 years ago
- CTF — 学习笔记&比赛题目&WP☆25Dec 9, 2023Updated 2 years ago
- Advanced Machine Learning Fall 2020 Project Repository☆12Dec 12, 2020Updated 5 years ago
- [EMNLP'22] Textual Manifold-based Defense Against Natural Language Adversarial Examples☆11Apr 6, 2023Updated 3 years ago
- [NeurIPS 2024] Fight Back Against Jailbreaking via Prompt Adversarial Tuning☆11Oct 29, 2024Updated last year
- Our research proposes a novel MoGU framework that improves LLMs' safety while preserving their usability.☆18Jan 14, 2025Updated last year
- Analyzing LLM Alignment via Token distribution shift☆17Jan 26, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Ultra-High-speed Spiking Recognition (UHSR)☆10May 20, 2024Updated 2 years ago
- Code Implementation of Adversarial Prompt Evaluation paper☆14Sep 18, 2025Updated 10 months ago
- Craft poisoned data using MetaPoison☆54Apr 5, 2021Updated 5 years ago
- The implementation of the Block Coordinate Regularization by Denoising (BC-RED) algorithm (NeurIPS 2019)☆10Oct 15, 2019Updated 6 years ago
- 百度AI安全对抗赛第一名团队示例代码,基于官方给出的PGD修改,主要内容为L2-PGD+EOT。☆11Mar 17, 2021Updated 5 years ago
- ☆20Jun 4, 2026Updated last month
- Blogs that I'm actively following.☆16Sep 17, 2023Updated 2 years ago
- Restore From Restored: Video Restoration With Pseudo Clean Video, CVPR 2021☆10Jun 20, 2021Updated 5 years ago
- Official repository of "Distort, Distract, Decode: Instruction-Tuned Model Can Refine its Response from Noisy Instructions", ICLR 2024 Sp…☆21Mar 7, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Code for NDSS '25 paper "Passive Inference Attacks on Split Learning via Adversarial Regularization"☆13Sep 16, 2024Updated last year
- [CVPRW'22] A privacy attack that exploits Adversarial Training models to compromise the privacy of Federated Learning systems.☆12Jul 7, 2022Updated 4 years ago
- ☆18Jul 2, 2023Updated 3 years ago
- NTIRE 2020 Real Image Denoising Challenge - ZJU231☆10Mar 26, 2020Updated 6 years ago
- ☆10Sep 20, 2023Updated 2 years ago
- UCAS大三自然语言处理课程大作业☆12Jun 25, 2023Updated 3 years ago
- 紫菜鱼的网络安全扫描器☆12Dec 19, 2023Updated 2 years ago