Official Codebase of the ACL 2026 Oral paper "Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring"
☆28Jun 25, 2026Updated 2 months ago
Alternatives and similar repositories for Jailbreak_Detection_RCS
Users that are interested in Jailbreak_Detection_RCS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026] Dissecting Failure Dynamics in Large Language Model Reasoning☆18Apr 17, 2026Updated 4 months ago
- MemoryDial☆15Mar 10, 2026Updated 6 months ago
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 10 months ago
- Data and Code Repository for “STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems”☆17Apr 17, 2026Updated 4 months ago
- ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing - ACL Findings 2026☆25Jul 15, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACL26 Findings] TopoDIM: One-shot Topology Generation of Diverse Interaction Modes for Multi-Agent Systems☆19Jan 19, 2026Updated 7 months ago
- [ACL2026] UCAS: Uncertainty-aware Advantage Shaping for RLVR☆32Apr 14, 2026Updated 5 months ago
- Implementation for paper Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs, which is accepted by ACL 2026 (main con…☆16Oct 10, 2025Updated 11 months ago
- [🏆CVPR'26] Official Repo for IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding☆34Jun 2, 2026Updated 3 months ago
- Curriculum-RLAIF is a data-centric curriculum learning framework for reward model training in RLAIF-based LLM alignment☆23Apr 18, 2026Updated 4 months ago
- This is the official repository for JailExpert☆23Sep 9, 2025Updated last year
- The code of the paper "Hybrid Relational Graphs with Sentiment-Laden Semantic Alignment for Multimodal Emotion Recognition in Conversatio…☆17Feb 22, 2026Updated 6 months ago
- videoPro: Adaptive Program Reasoning for Long Video Understanding☆45Apr 15, 2026Updated 5 months ago
- ACL26 Long Paper☆19Jul 4, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Sparse Adapter Fusion for Continual Learning in NLP - EACL 2026☆16Apr 9, 2026Updated 5 months ago
- [Paper][EMNLP 2025] Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching☆35Feb 8, 2026Updated 7 months ago
- ☆79Apr 12, 2026Updated 5 months ago
- This is the official repo for the paper "General365: Benchmarking General Reasoning in LLMs under High Difficulty and Diversity".☆89Apr 14, 2026Updated 5 months ago
- [ACL 2026] Context-Agent: Dynamic Discourse Trees for Non-Linear Dialogue☆24Apr 14, 2026Updated 5 months ago
- This is the official repo for the paper "AMO-Bench: Large Language Models Still Struggle in High School Math Competitions".☆172Feb 6, 2026Updated 7 months ago
- [AAAI2024] Debiasing Multimodal Sarcasm Detection with Contrastive Learning☆17Jan 5, 2024Updated 2 years ago
- [COLM 2025] JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model☆26Nov 25, 2025Updated 9 months ago
- [ICLR 2025] PyTorch Implementation of "ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time"☆34Jul 20, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆43Jul 3, 2026Updated 2 months ago
- SE-Agent is a self-evolution framework for LLM Code agents. It enables trajectory-level evolution to exchange information across reasonin…☆286Sep 23, 2025Updated 11 months ago
- [CVPR Findings 2026] HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model☆17Mar 8, 2026Updated 6 months ago
- ☆16Jul 2, 2026Updated 2 months ago
- 📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for str…☆204Updated this week
- [COLM 2024] JailBreakV-28K: A comprehensive benchmark designed to evaluate the transferability of LLM jailbreak attacks to MLLMs, and fur …☆97May 9, 2025Updated last year
- A curated collection of research and techniques for protecting intellectual property of large language models, including watermarking, fi…☆53Jun 10, 2026Updated 3 months ago
- ☆13Jan 14, 2025Updated last year
- Assignments of CSCE-642: Deep Reinforcement Learning offered at Texas A&M University.☆10Aug 31, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆71Jun 1, 2025Updated last year
- ☆46Jun 28, 2025Updated last year
- A framework for steering MoE models by detecting and controlling behavior-linked experts.☆38Sep 12, 2025Updated last year
- ☆14Jul 28, 2025Updated last year
- [CVPR 2025] Official implementation for "Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbre…☆63Jul 5, 2025Updated last year
- [CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment☆29Jun 11, 2025Updated last year
- code for infocom 2021 paper MANDA☆11May 30, 2023Updated 3 years ago