Code for paper "MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM safety"
☆58Oct 1, 2026Updated last week
Alternatives and similar repositories for MAGIC
Users that are interested in MAGIC are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TRACE, a framework for turn-aware credit assignment for multi-turn jailbreak optimization☆22Aug 26, 2026Updated last month
- [🏆CVPR'26] Official Repo for IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding☆34Sep 17, 2026Updated 3 weeks ago
- ☆33Apr 9, 2026Updated 6 months ago
- Build your agent from 200,000+ skills via skill RETRIEVAL & ORCHESTRATION☆617Mar 7, 2026Updated 7 months ago
- [CVPR 2026] SOTA Chemical Reaction Diagram Parsing Framework☆29Mar 24, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [Pattern Recognition 2025 🌟]Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation☆11Jun 12, 2024Updated 2 years ago
- [DAI 2025] Beyond GPT-5: Making LLMs Cheaper and Better via Performance–Efficiency Optimized Routing☆223Dec 11, 2025Updated 9 months ago
- Official implementation of "TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards"☆32Sep 29, 2026Updated last week
- A Diagnostic Guardrail Framework for AI Agent Safety and Security☆701Jun 8, 2026Updated 4 months ago
- An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation☆61Jun 4, 2026Updated 4 months ago
- DARWIN is a self-evolving LLM jailbreak framework that grows a reusable strategy pool through external extraction, sandbox filtering, his…☆80Sep 28, 2026Updated last week
- ☆172Mar 20, 2026Updated 6 months ago
- offical implementation of MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming☆18Jun 2, 2025Updated last year
- Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation☆16Mar 28, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The official implementation of A Unified Game-Theoretic Interpretation of Adversarial Robustness.☆22Jun 9, 2022Updated 4 years ago
- A Framework for Evaluating AI Agent Safety in Realistic Environments☆38Aug 10, 2026Updated 2 months ago
- The Code for the EMNLP 2023 main conference paper "Prompt-based Logical Semantics Enhancement for Implicit Discourse Relation Recognition…☆13Dec 10, 2023Updated 2 years ago
- [AAAI 2026] The official code for ``LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation''☆17Mar 20, 2026Updated 6 months ago
- [Findings@ACL'26] LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing☆118Apr 6, 2026Updated 6 months ago
- Catalogue of Life toolkit for Python☆11Aug 4, 2020Updated 6 years ago
- This repo is the official implementation of “Are Your Agents Upward Deceivers?”. The paper is accepted by ICML 2026.☆24Dec 15, 2025Updated 9 months ago
- Universal preflight security scanner for AI coding agents — Detects hooks injection, credential exfiltration & backdoors in .cursorrules,…☆77May 29, 2026Updated 4 months ago
- Code of paper: xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking"☆18Apr 3, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆48May 29, 2026Updated 4 months ago
- ☆18Mar 25, 2026Updated 6 months ago
- ☆26Mar 17, 2025Updated last year
- ☆24Nov 19, 2024Updated last year
- PIForge: An Open Framework for RL-based Prompt Injection Red Teaming.☆35Sep 30, 2026Updated last week
- [ArXiv 2025] Imperceptible Jailbreaking against Large Language Models☆25Oct 7, 2025Updated last year
- Paper list for LLM/MLLM-based image segmentation☆49Dec 24, 2025Updated 9 months ago
- Implementation of SBM-meet-GNN☆23May 12, 2019Updated 7 years ago
- Code for generative hypergraph clustering via modularity-like objective functions.☆28Jun 10, 2022Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated last year
- We are dedicated to building a set of open agent skills that deliver superior performance, higher determinism, and greater consistency on…☆133Dec 29, 2025Updated 9 months ago
- The implementation for paper "UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in …☆17Jul 3, 2025Updated last year
- Official code for "TWINS: A Fine-Tuning Framework for Improved Transferability of Adversarial Robustness and Generalization", CVPR 2023☆13Apr 26, 2023Updated 3 years ago
- All-in-One Safety Evaluation Framwork☆56Aug 12, 2026Updated last month
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆20Oct 3, 2026Updated last week
- Solving the OpenAI Gym (MountainCarContinuous-v0) with DDPG☆21Jan 23, 2023Updated 3 years ago