Code for paper "MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM safety"
☆57Aug 26, 2026Updated 3 weeks ago
Alternatives and similar repositories for MAGIC
Users that are interested in MAGIC are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆31May 14, 2026Updated 4 months ago
- [🏆CVPR'26] Official Repo for IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding☆34Updated this week
- ☆32Apr 9, 2026Updated 5 months ago
- Velaclaw — control plane for team AI☆110Jul 21, 2026Updated last month
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision☆30May 26, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Open-source red teaming framework for MLLMs with 42+ attack methods☆272Jul 17, 2026Updated 2 months ago
- A Diagnostic Guardrail Framework for AI Agent Safety and Security☆700Jun 8, 2026Updated 3 months ago
- DARWIN is a self-evolving LLM jailbreak framework that grows a reusable strategy pool through external extraction, sandbox filtering, his…☆74Jun 5, 2026Updated 3 months ago
- offical implementation of MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming☆18Jun 2, 2025Updated last year
- ☆45Oct 21, 2025Updated 10 months ago
- The official implementation of A Unified Game-Theoretic Interpretation of Adversarial Robustness.☆22Jun 9, 2022Updated 4 years ago
- A Framework for Evaluating AI Agent Safety in Realistic Environments☆38Aug 10, 2026Updated last month
- [ICLR 2026] TwinVLA : Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models☆22May 29, 2026Updated 3 months ago
- The Code for the EMNLP 2023 main conference paper "Prompt-based Logical Semantics Enhancement for Implicit Discourse Relation Recognition…☆13Dec 10, 2023Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [AAAI 2026] The official code for ``LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation''☆17Mar 20, 2026Updated 6 months ago
- Universal preflight security scanner for AI coding agents — Detects hooks injection, credential exfiltration & backdoors in .cursorrules,…☆76May 29, 2026Updated 3 months ago
- Code of paper: xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking"☆17Apr 3, 2026Updated 5 months ago
- ☆18Mar 25, 2026Updated 5 months ago
- ☆27Mar 17, 2025Updated last year
- ☆24Nov 19, 2024Updated last year
- PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses(COLM 2026)☆28Aug 27, 2026Updated 3 weeks ago
- Code accompanying paper "Coordinated Proximal Policy Optimization"☆10Mar 26, 2022Updated 4 years ago
- [ICLR'26 Oral] RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments☆64Feb 9, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ArXiv 2025] Imperceptible Jailbreaking against Large Language Models☆25Oct 7, 2025Updated 11 months ago
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated last year
- The implementation for paper "UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in …☆17Jul 3, 2025Updated last year
- Official code for "TWINS: A Fine-Tuning Framework for Improved Transferability of Adversarial Robustness and Generalization", CVPR 2023☆13Apr 26, 2023Updated 3 years ago
- ☆17Nov 3, 2024Updated last year
- ☆17Sep 3, 2025Updated last year
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆19Updated this week
- Implementation for paper Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs, which is accepted by ACL 2026 (main con…☆16Oct 10, 2025Updated 11 months ago
- Solving the OpenAI Gym (MountainCarContinuous-v0) with DDPG☆21Jan 23, 2023Updated 3 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Agent Security Bench (ASB)☆300Apr 16, 2026Updated 5 months ago
- All Claude Code Prompts☆15Apr 2, 2026Updated 5 months ago
- MinT-2M: Long-context training system for resident-prefix GRPO☆46Jul 24, 2026Updated last month
- ☆20Sep 17, 2025Updated last year
- Online Preference Alignment for Language Models via Count-based Exploration☆21Jan 14, 2025Updated last year
- [ACL 2024] CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion☆60Oct 1, 2025Updated 11 months ago
- [AAAI 2023 Oral] Official code for "PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction".☆21Jul 26, 2025Updated last year