S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
☆120Feb 13, 2026Updated 6 months ago
Alternatives and similar repositories for S-Eval
Users that are interested in S-Eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆21May 31, 2024Updated 2 years ago
- Flames is a highly adversarial benchmark in Chinese for LLM's harmlessness evaluation developed by Shanghai AI Lab and Fudan NLP Group.☆68May 21, 2024Updated 2 years ago
- SC-Safety: 中文大模型多轮对抗安全基准☆152Mar 15, 2024Updated 2 years ago
- The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-hous…☆22Sep 3, 2025Updated 11 months ago
- A curated list of awesome publications and researchers on prompting framework updated and maintained by The Intelligent System Security (…☆88Jan 14, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆10Mar 13, 2023Updated 3 years ago
- 【ACL 2024】 SALAD benchmark & MD-Judge☆175Mar 8, 2025Updated last year
- Official repo for GPTFUZZER : Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts☆606Feb 27, 2026Updated 6 months ago
- Code release for RobOT (ICSE'21)☆15Dec 5, 2022Updated 3 years ago
- Code release for DeepJudge (S&P'22)☆52Mar 14, 2023Updated 3 years ago
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- ☆43Dec 8, 2025Updated 8 months ago
- ☆27Feb 1, 2023Updated 3 years ago
- White-box Fairness Testing through Adversarial Sampling☆14Apr 16, 2021Updated 5 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The official code for ``An Engorgio Prompt Makes Large Language Model Babble on''☆24Aug 9, 2025Updated last year
- AISafetyLab: A comprehensive framework covering safety attack, defense, evaluation and paper list.☆250Apr 21, 2026Updated 4 months ago
- ☆31Oct 23, 2024Updated last year
- "他山之石、可以攻玉":复旦JADE团队发布的大模型测评与治理系列☆526Jul 23, 2026Updated last month
- Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。☆1,215Feb 27, 2024Updated 2 years ago
- ☆23Jan 14, 2025Updated last year
- The code implementation of MuScleLoRA (Accepted in ACL 2024)☆10Dec 1, 2024Updated last year
- SuperCLUE-Agent: 基于中文原生任务的Agent智能体核心能力测评基准☆96Nov 9, 2023Updated 2 years ago
- The first multi-level safety evaluation platform for OpenClaw-style AI agents.☆30May 27, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs☆343Jun 7, 2024Updated 2 years ago
- Implementation of TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems (https://arxiv.org/pdf/190…☆19Apr 13, 2023Updated 3 years ago
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"☆10Jun 23, 2024Updated 2 years ago
- ☆17Jun 27, 2021Updated 5 years ago
- HOD: A Benchmark Dataset for Harmful Object Detection☆38Jun 11, 2025Updated last year
- Accepted by ECCV 2024☆218Oct 15, 2024Updated last year
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- ☆29Mar 10, 2026Updated 5 months ago
- [EMNLP 2025] Official code for the paper "SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning"☆16May 12, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆52Feb 25, 2026Updated 6 months ago
- Code for paper: AdvKnn: Adversarial Attacks On K-Nearest Neighbor Classifiers With Approximate Gradients☆14Dec 23, 2019Updated 6 years ago
- ☆11Jan 3, 2024Updated 2 years ago
- Advances in Neural Information Processing Systems (NeurIPS 2021)☆23Nov 4, 2022Updated 3 years ago
- ☆37Jan 7, 2025Updated last year
- [ICLR 2024]Data for "Multilingual Jailbreak Challenges in Large Language Models"☆108Mar 7, 2024Updated 2 years ago
- A curated list of safety-related papers, articles, and resources focused on Large Language Models (LLMs). This repository aims to provide…☆1,901Jul 12, 2026Updated last month