Official github repo for SafeDialBench, a comprehensive multi-turn dialogue benchmark to evaluate LLMs' safety.
☆56May 12, 2025Updated last year
Alternatives and similar repositories for SafeDialBench-Dataset
Users that are interested in SafeDialBench-Dataset are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning☆23Jun 23, 2026Updated last month
- ☆15May 27, 2025Updated last year
- Implementation of the paper "Multi-Agent Exploration via Self-Learning and Social Learning"☆20Dec 7, 2024Updated last year
- A Framework of Continual Learning☆136Dec 9, 2025Updated 8 months ago
- Implementation of the paper "WToE: Learning When to Explore in Multi-Agent Reinforcement Learning"☆21Aug 17, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".☆19Jun 24, 2026Updated last month
- Implementation of the paper "Egoism, Utilitarianism and Egalitarianism in Multi-Agent Reinforcement Learning"☆21Aug 17, 2024Updated 2 years ago
- Implementation of the paper "Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in Mixed Coo…☆17Dec 7, 2024Updated last year
- [ICCV 2025] Official code of paper "Dynamic Multi-Layer Null Space Projection for Vision-Language Continual Learning"☆27Sep 8, 2025Updated 11 months ago
- ManifoldAlignmentStyleTransfer☆46Feb 24, 2022Updated 4 years ago
- Red Queen Dataset and data generation template☆27Dec 26, 2025Updated 7 months ago
- offical implementation of MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming☆18Jun 2, 2025Updated last year
- ☆21Apr 7, 2025Updated last year
- Unpaired Caricature Generation with Multiple Exaggerations (TMM 2021)☆40Jul 14, 2021Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types☆26Nov 29, 2024Updated last year
- ☆142Dec 3, 2025Updated 8 months ago
- ☆67May 21, 2025Updated last year
- Platform for training generalizable deep reinforcement learning agents☆15Mar 4, 2026Updated 5 months ago
- ☆134Jul 29, 2026Updated 2 weeks ago
- Official PyTorch implementation of "ACE:Off-Policy Actor-Critic with Causality-Aware Entropy Regularization"☆35May 13, 2024Updated 2 years ago
- Official eval code for ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation☆27Dec 12, 2025Updated 8 months ago
- ☆23Jun 18, 2025Updated last year
- Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency☆16Aug 6, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official Implementation of "ToolSafe: Enhancing Tool Invocation Safety of LLM-based Agents via Proactive Step-level Guardrail and Feedbac…☆75Mar 25, 2026Updated 4 months ago
- Exploiting Inter-sample and Inter-feature Relations in Dataset Distillation (CVPR24)☆10Jun 16, 2024Updated 2 years ago
- Code for the paper: Causal Action Influence Aware Counterfactual Data Augmentation @ICML2024☆13Jul 19, 2024Updated 2 years ago
- [ICML 2026] Official implementations of ``SafeSearch: Automated Red-Teaming of LLM-Based Search Agents''☆19Mar 25, 2026Updated 4 months ago
- [COLM 2025] JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model☆26Nov 25, 2025Updated 8 months ago
- Source Code for "Adapters for Enhanced Modeling of Multilingual Knowledge and Text"☆12Oct 28, 2022Updated 3 years ago
- ☆17Dec 23, 2024Updated last year
- Code for "LifeLong Incremental Reinforcement Learning (LLIRL)"☆21Jan 28, 2021Updated 5 years ago
- ☆11Oct 9, 2022Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Bird’s Eye: Probing for Linguistic Graph Structureswith a Simple Information-Theoretic Approach☆11Aug 1, 2021Updated 5 years ago
- [NeurIPS 2025]: Personalized Safety in LLMs — A Benchmark and a Planning-Based Agent Approach☆18Oct 30, 2025Updated 9 months ago
- [ACL 2024] MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues☆154Jul 24, 2024Updated 2 years ago
- Official code for paper "Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning"☆15Jun 12, 2025Updated last year
- knowledge distillation for few-shot learning☆13Dec 27, 2023Updated 2 years ago
- The paper list of the 86-page paper "The Rise and Potential of Large Language Model Based Agents: A Survey" by Zhiheng Xi et al.☆12May 2, 2024Updated 2 years ago
- Adversarial attacks including DeepFool and C&W☆13May 20, 2019Updated 7 years ago