[COLM 2025] JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
☆26Nov 25, 2025Updated 9 months ago
Alternatives and similar repositories for JailDAM
Users that are interested in JailDAM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is the official repository for Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language Models.☆16Jan 16, 2025Updated last year
- ☆72Sep 30, 2025Updated 11 months ago
- Official code of "The Automated but Risky Game: Modeling and Benchmarking Agent-to-Agent Negotiations and Transactions in Consumer Market…☆31Jun 9, 2026Updated 2 months ago
- ☆83Mar 30, 2025Updated last year
- [CVPR 2025] Official implementation for JOOD "Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy"☆22Jun 11, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ACL 2025 Findings] Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducemen…☆37Jun 29, 2025Updated last year
- [ICLR 2025] PyTorch Implementation of "ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time"☆34Jul 20, 2025Updated last year
- Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency☆16Aug 6, 2025Updated last year
- [CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment☆29Jun 11, 2025Updated last year
- [COLM 2024] JailBreakV-28K: A comprehensive benchmark designed to evaluate the transferability of LLM jailbreak attacks to MLLMs, and fur…☆97May 9, 2025Updated last year
- ☆46Jun 19, 2025Updated last year
- [ACL 2025] Data and Code for Paper VLSBench: Unveiling Visual Leakage in Multimodal Safety☆62Jul 21, 2025Updated last year
- ☆62Jun 5, 2024Updated 2 years ago
- Official code for the paper: "Multi-User Large Language Model Agents"☆34Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [ICLR 2025] Official codebase for the ICLR 2025 paper "Multimodal Situational Safety"☆37Jun 23, 2025Updated last year
- Official Codebase of the ACL 2026 Oral paper "Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contra…☆28Jun 25, 2026Updated 2 months ago
- ACL 2025 (Main) HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States☆166Jun 8, 2025Updated last year
- ☆16Aug 30, 2025Updated last year
- ☆12Nov 21, 2023Updated 2 years ago
- Röttger et al. (2025): "MSTS: A Multimodal Safety Test Suite for Vision-Language Models"☆20Mar 31, 2025Updated last year
- 😎 up-to-date & curated list of awesome Attacks on Large-Vision-Language-Models papers, methods & resources.☆575Updated this week
- ☆40May 17, 2025Updated last year
- [CVPR 2026] LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models☆16Aug 30, 2026Updated last week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [ ICCV'2025 Poster ] SecDOOD is a secure cloud-device collaboration framework for efficient on-device OOD detection without requiring de…☆15Oct 10, 2025Updated 10 months ago
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools☆16Dec 19, 2025Updated 8 months ago
- ☆67May 21, 2025Updated last year
- [TMLR'25] AutoTrust, a groundbreaking benchmark designed to assess the trustworthiness of DriveVLMs. This work aims to enhance public saf…☆54Nov 20, 2025Updated 9 months ago
- Code for ACM MM2024 paper: White-box Multimodal Jailbreaks Against Large Vision-Language Models☆34Dec 30, 2024Updated last year
- [ICML 2024] Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.☆91Jan 19, 2025Updated last year
- ☆42Jul 3, 2026Updated 2 months ago
- ☆21Apr 7, 2025Updated last year
- Code of paper: xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking"☆17Apr 3, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".☆19Jun 24, 2026Updated 2 months ago
- Code repo of our paper Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis (https://arxiv.org/abs/2406.10794…☆24Jul 26, 2024Updated 2 years ago
- The code for the paper "Pre-trained Vision-Language Models Learn Discoverable Concepts"☆21Jun 5, 2024Updated 2 years ago
- [ACM MM 2024] ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack☆14Dec 20, 2024Updated last year
- A curated list of resources dedicated to the safety of Large Vision-Language Models. This repository aligns with our survey titled A Surv…☆218May 25, 2026Updated 3 months ago
- Official codebase for "STAIR: Improving Safety Alignment with Introspective Reasoning"☆89Feb 26, 2025Updated last year
- ☆24Aug 28, 2026Updated last week