Official Codebase of the ACL 2026 Oral paper "Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring"
☆26Jun 25, 2026Updated last month
Alternatives and similar repositories for Jailbreak_Detection_RCS
Users that are interested in Jailbreak_Detection_RCS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026] Dissecting Failure Dynamics in Large Language Model Reasoning☆18Apr 17, 2026Updated 3 months ago
- MemoryDial☆15Mar 10, 2026Updated 4 months ago
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 9 months ago
- Data and Code Repository for “STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems”☆17Apr 17, 2026Updated 3 months ago
- ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing - ACL Findings 2026☆25Jul 15, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACL26 Findings] TopoDIM: One-shot Topology Generation of Diverse Interaction Modes for Multi-Agent Systems☆19Jan 19, 2026Updated 6 months ago
- [ACL2026] UCAS: Uncertainty-aware Advantage Shaping for RLVR☆31Apr 14, 2026Updated 3 months ago
- Implementation for paper Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs, which is accepted by ACL 2026 (main con…☆16Oct 10, 2025Updated 9 months ago
- [🏆CVPR'26] Official Repo for IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding☆33Jun 2, 2026Updated 2 months ago
- Curriculum-RLAIF is a data-centric curriculum learning framework for reward model training in RLAIF-based LLM alignment☆23Apr 18, 2026Updated 3 months ago
- This is the official repository for JailExpert☆23Sep 9, 2025Updated 10 months ago
- Prototype Conditioned Generative Replay for Continual Learning in NLP - NAACL 2025☆26Updated this week
- ☆17Feb 22, 2026Updated 5 months ago
- videoPro: Adaptive Program Reasoning for Long Video Understanding☆45Apr 15, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ACL26 Long Paper☆18Jul 4, 2026Updated last month
- Sparse Adapter Fusion for Continual Learning in NLP - EACL 2026☆15Apr 9, 2026Updated 3 months ago
- [Paper][EMNLP 2025] Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching☆35Feb 8, 2026Updated 5 months ago
- ☆79Apr 12, 2026Updated 3 months ago
- This is the official repo for the paper "General365: Benchmarking General Reasoning in LLMs under High Difficulty and Diversity".☆88Apr 14, 2026Updated 3 months ago
- [🏆AAAI'25] Official Repo for ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area.☆88Apr 14, 2026Updated 3 months ago
- This is the official repo for the paper "AMO-Bench: Large Language Models Still Struggle in High School Math Competitions".☆177Feb 6, 2026Updated 5 months ago
- The implementation of ACL 2026 paper "Rethinking entropy interventions in rlvr: An entropy change perspective"☆26Jul 19, 2026Updated 2 weeks ago
- [COLM 2025] JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model☆26Nov 25, 2025Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆18Jul 1, 2026Updated last month
- [ICLR 2025] PyTorch Implementation of "ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time"☆34Jul 20, 2025Updated last year
- ☆41Jul 3, 2026Updated last month
- SE-Agent is a self-evolution framework for LLM Code agents. It enables trajectory-level evolution to exchange information across reasonin…☆281Sep 23, 2025Updated 10 months ago
- [CVPR Findings 2026] HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model☆17Mar 8, 2026Updated 4 months ago
- ☆12Jul 2, 2026Updated last month
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆12Oct 31, 2024Updated last year
- 📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for str…☆191Jul 22, 2026Updated last week
- [AAAI 2026] This is the official implementation of the paper "ExtendAttack: Attacking Servers of LRMs via Extending Reasoning".☆25Mar 18, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [NeurIPS 2025] The official implementation of the paper "DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agen…☆59Jul 16, 2026Updated 2 weeks ago
- [COLM 2024] JailBreakV-28K: A comprehensive benchmark designed to evaluate the transferability of LLM jailbreak attacks to MLLMs, and fur…☆96May 9, 2025Updated last year
- A curated collection of research and techniques for protecting intellectual property of large language models, including watermarking, fi…☆52Jun 10, 2026Updated last month
- ☆13Jan 14, 2025Updated last year
- ☆70Jun 1, 2025Updated last year
- ☆43Jun 28, 2025Updated last year
- ☆14Jul 28, 2025Updated last year