Official github repo for SafeDialBench, a comprehensive multi-turn dialogue benchmark to evaluate LLMs' safety.
☆57May 12, 2025Updated last year
Alternatives and similar repositories for SafeDialBench-Dataset
Users that are interested in SafeDialBench-Dataset are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning☆23Jun 23, 2026Updated 3 months ago
- A Framework of Continual Learning☆135Dec 9, 2025Updated 9 months ago
- Implementation of the paper "Egoism, Utilitarianism and Egalitarianism in Multi-Agent Reinforcement Learning"☆21Aug 17, 2024Updated 2 years ago
- Implementation of the paper "Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in Mixed Coo…☆17Dec 7, 2024Updated last year
- ☆27May 14, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICCV 2025] Official code of paper "Dynamic Multi-Layer Null Space Projection for Vision-Language Continual Learning"☆28Sep 8, 2025Updated last year
- ManifoldAlignmentStyleTransfer☆46Feb 24, 2022Updated 4 years ago
- Red Queen Dataset and data generation template☆29Dec 26, 2025Updated 9 months ago
- offical implementation of MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming☆18Jun 2, 2025Updated last year
- ☆21Apr 7, 2025Updated last year
- Unpaired Caricature Generation with Multiple Exaggerations (TMM 2021)☆40Jul 14, 2021Updated 5 years ago
- [TPAMI 2023] LibFewShot: A Comprehensive Library for Few-shot Learning.☆1,072Oct 27, 2025Updated 11 months ago
- SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types☆26Nov 29, 2024Updated last year
- ☆145Dec 3, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆69May 21, 2025Updated last year
- ☆134Jul 29, 2026Updated last month
- ☆23Jun 18, 2025Updated last year
- Unofficial Pytorch(1.0+) implementation of ICCV 2019 paper "Multimodal Style Transfer via Graph Cuts"☆16Jan 9, 2020Updated 6 years ago
- Official eval code for ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation☆29Dec 12, 2025Updated 9 months ago
- Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency☆16Aug 6, 2025Updated last year
- Exploiting Inter-sample and Inter-feature Relations in Dataset Distillation (CVPR24)☆10Jun 16, 2024Updated 2 years ago
- Code for the paper: Causal Action Influence Aware Counterfactual Data Augmentation @ICML2024☆14Jul 19, 2024Updated 2 years ago
- ☆15Jun 2, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆23Jul 26, 2025Updated last year
- [ICML 2026] Official implementations of ``SafeSearch: Automated Red-Teaming of LLM-Based Search Agents''☆19Mar 25, 2026Updated 6 months ago
- 南京大学 NJU 计算机系 CS 课程资料 作业 代码 实验报告(数据挖掘 模式识别 机器学习导论 概率论与数理统计 计算机图形学 高级程序设计 数据库 计算机系统基础 操作系统 程设实验 数电 数电实验... ) 更新中, star!☆23Jun 28, 2020Updated 6 years ago
- ☆91Jun 19, 2026Updated 3 months ago
- [NeurIPS 2025]: Personalized Safety in LLMs — A Benchmark and a Planning-Based Agent Approach☆18Oct 30, 2025Updated 10 months ago
- [ACL 2024] MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues☆153Jul 24, 2024Updated 2 years ago
- Official code for paper "Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning"☆16Jun 12, 2025Updated last year
- Official PyTorch code for "Sample Efficient Offline-to-Online Reinforcement Learning" in TKDE'23.☆16Aug 14, 2023Updated 3 years ago
- knowledge distillation for few-shot learning☆13Dec 27, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The paper list of the 86-page paper "The Rise and Potential of Large Language Model Based Agents: A Survey" by Zhiheng Xi et al.☆12May 2, 2024Updated 2 years ago
- Adversarial attacks including DeepFool and C&W☆13May 20, 2019Updated 7 years ago
- Design for Error Detection in Deep-Research Agents Trajectories.☆24Jun 4, 2026Updated 3 months ago
- Data and code for the paper: Finding Safety Neurons in Large Language Models☆31Jan 29, 2026Updated 7 months ago
- ☆17Oct 11, 2022Updated 3 years ago
- Official codebase for "STAIR: Improving Safety Alignment with Introspective Reasoning"☆89Feb 26, 2025Updated last year
- The code of "Deep Embedded Complementary and Interactive Information for Multi-view Classification", AAAI 2020.☆12May 28, 2020Updated 6 years ago