S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
☆121Feb 13, 2026Updated 7 months ago
Alternatives and similar repositories for S-Eval
Users that are interested in S-Eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flames is a highly adversarial benchmark in Chinese for LLM's harmlessness evaluation developed by Shanghai AI Lab and Fudan NLP Group.☆68May 21, 2024Updated 2 years ago
- SC-Safety: 中文大模型多轮对抗安全基准☆152Mar 15, 2024Updated 2 years ago
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated last year
- ☆10Mar 13, 2023Updated 3 years ago
- Official github repo for SafetyBench, a comprehensive benchmark to evaluate LLMs' safety. [ACL 2024]☆296Jul 28, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 【ACL 2024】 SALAD benchmark & MD-Judge☆175Mar 8, 2025Updated last year
- Code release for RobOT (ICSE'21)☆15Dec 5, 2022Updated 3 years ago
- ☆16Apr 23, 2025Updated last year
- Code release for DeepJudge (S&P'22)☆52Mar 14, 2023Updated 3 years ago
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- ☆27Feb 1, 2023Updated 3 years ago
- White-box Fairness Testing through Adversarial Sampling☆14Apr 16, 2021Updated 5 years ago
- The official code for ``An Engorgio Prompt Makes Large Language Model Babble on''☆24Aug 9, 2025Updated last year
- AISafetyLab: A comprehensive framework covering safety attack, defense, evaluation and paper list.☆251Apr 21, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆33Oct 23, 2024Updated last year
- The repo for using the model https://huggingface.co/thu-coai/Attacker-v0.1☆13Apr 23, 2025Updated last year
- Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。☆1,218Feb 27, 2024Updated 2 years ago
- ☆23Jan 14, 2025Updated last year
- The code implementation of MuScleLoRA (Accepted in ACL 2024)☆11Dec 1, 2024Updated last year
- SuperCLUE-Agent: 基于中文原生任务的Agent智能体核心能力测评基准☆96Nov 9, 2023Updated 2 years ago
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs☆344Jun 7, 2024Updated 2 years ago
- Implementation of TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems (https://arxiv.org/pdf/190…☆19Apr 13, 2023Updated 3 years ago
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"☆10Jun 23, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- HOD: A Benchmark Dataset for Harmful Object Detection☆38Jun 11, 2025Updated last year
- Accepted by ECCV 2024☆221Oct 15, 2024Updated last year
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- ☆29Mar 10, 2026Updated 6 months ago
- [EMNLP 2025] Official code for the paper "SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning"☆16May 12, 2026Updated 4 months ago
- ☆53Feb 25, 2026Updated 6 months ago
- The first multi-level safety evaluation platform for OpenClaw-style AI agents.☆37May 27, 2026Updated 3 months ago
- Code for paper: AdvKnn: Adversarial Attacks On K-Nearest Neighbor Classifiers With Approximate Gradients☆14Dec 23, 2019Updated 6 years ago
- ☆11Jan 3, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Advances in Neural Information Processing Systems (NeurIPS 2021)☆23Nov 4, 2022Updated 3 years ago
- ☆37Jan 7, 2025Updated last year
- A curated list of safety-related papers, articles, and resources focused on Large Language Models (LLMs). This repository aims to provide…☆1,911Jul 12, 2026Updated 2 months ago
- Code to enable layer-level steering in LLMs using sparse auto encoders☆35Sep 18, 2025Updated last year
- ☆29Jun 5, 2024Updated 2 years ago
- Universal and Transferable Attacks on Aligned Language Models☆4,800Aug 2, 2024Updated 2 years ago
- A simple evaluation of generative language models and safety classifiers.☆109Aug 31, 2026Updated 2 weeks ago