Official github repo for SafeDialBench, a comprehensive multi-turn dialogue benchmark to evaluate LLMs' safety.
☆57May 12, 2025Updated last year
Alternatives and similar repositories for SafeDialBench-Dataset
Users that are interested in SafeDialBench-Dataset are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of the paper "Multi-Agent Exploration via Self-Learning and Social Learning"☆20Dec 7, 2024Updated last year
- RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation☆18Oct 16, 2025Updated 10 months ago
- A Framework of Continual Learning☆134Dec 9, 2025Updated 8 months ago
- Implementation of the paper "WToE: Learning When to Explore in Multi-Agent Reinforcement Learning"☆21Aug 17, 2024Updated 2 years ago
- The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".☆19Jun 24, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementation of the paper "Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in Mixed Coo…☆17Dec 7, 2024Updated last year
- ☆24May 14, 2025Updated last year
- Red Queen Dataset and data generation template☆29Dec 26, 2025Updated 8 months ago
- offical implementation of MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming☆18Jun 2, 2025Updated last year
- [TPAMI 2023] LibFewShot: A Comprehensive Library for Few-shot Learning.☆1,069Oct 27, 2025Updated 10 months ago
- ☆142Dec 3, 2025Updated 9 months ago
- ☆67May 21, 2025Updated last year
- ☆132Jul 29, 2026Updated last month
- Official PyTorch implementation of "ACE:Off-Policy Actor-Critic with Causality-Aware Entropy Regularization"☆35May 13, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Open benchmark for AI agent security tools — prompt injection, data exfiltration, tool abuse, provenance☆25Aug 26, 2026Updated last week
- ☆23Jun 18, 2025Updated last year
- Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency☆16Aug 6, 2025Updated last year
- Official Implementation of "ToolSafe: Enhancing Tool Invocation Safety of LLM-based Agents via Proactive Step-level Guardrail and Feedbac…☆78Mar 25, 2026Updated 5 months ago
- ☆14Nov 12, 2024Updated last year
- This is the Pytorch implementation of paper--Training deep neural-networks using a noise adaptation layer.☆10Apr 18, 2021Updated 5 years ago
- ☆15Jun 2, 2026Updated 3 months ago
- ☆23Jul 26, 2025Updated last year
- [ICML 2026] Official implementations of ``SafeSearch: Automated Red-Teaming of LLM-Based Search Agents''☆19Mar 25, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 南京大学 NJU 计算机系 CS 课程资料 作业 代码 实验报告(数据挖掘 模式识别 机器学习导论 概率论与数理统计 计算机图形学 高级程序设计 数据库 计算机系统基础 操作系统 程设实验 数电 数电实验... ) 更新中, star!☆22Jun 28, 2020Updated 6 years ago
- ☆17Dec 23, 2024Updated last year
- Code for "LifeLong Incremental Reinforcement Learning (LLIRL)"☆21Jan 28, 2021Updated 5 years ago
- Bird’s Eye: Probing for Linguistic Graph Structureswith a Simple Information-Theoretic Approach☆11Aug 1, 2021Updated 5 years ago
- [NeurIPS 2025]: Personalized Safety in LLMs — A Benchmark and a Planning-Based Agent Approach☆18Oct 30, 2025Updated 10 months ago
- [ACL 2024] MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues☆154Jul 24, 2024Updated 2 years ago
- Official code for paper "Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning"☆16Jun 12, 2025Updated last year
- Official PyTorch code for "Sample Efficient Offline-to-Online Reinforcement Learning" in TKDE'23.☆16Aug 14, 2023Updated 3 years ago
- The Pytorch code of "Asymmetric Distribution Measure for Few-shot Learning", IJCAI 2020.☆15Oct 9, 2020Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The paper list of the 86-page paper "The Rise and Potential of Large Language Model Based Agents: A Survey" by Zhiheng Xi et al.☆12May 2, 2024Updated 2 years ago
- This repository contains a collection of the most influential papers, and benchmarks related to Large Language Models (LLMs) based Agent …☆59Jul 7, 2025Updated last year
- Design for Error Detection in Deep-Research Agents Trajectories.☆23Jun 4, 2026Updated 3 months ago
- ☆18Jan 6, 2025Updated last year
- Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models☆16Nov 4, 2023Updated 2 years ago
- Data and code for the paper: Finding Safety Neurons in Large Language Models☆31Jan 29, 2026Updated 7 months ago
- Learning Safety Constraints for Large Language Models (ICML2025)☆35May 25, 2026Updated 3 months ago