A simple evaluation of generative language models and safety classifiers.
☆105Jun 16, 2026Updated last month
Alternatives and similar repositories for safety-eval
Users that are interested in safety-eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An official codebase for "NormLens: Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Comm…☆10May 9, 2024Updated 2 years ago
- Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs☆131Dec 2, 2024Updated last year
- ☆17Mar 10, 2026Updated 4 months ago
- ☆42Aug 10, 2024Updated last year
- SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types☆26Nov 29, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Corpus to accompany: "Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding"☆11Apr 11, 2025Updated last year
- 【ACL 2024】 SALAD benchmark & MD-Judge☆176Mar 8, 2025Updated last year
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- Learning to route instances for Human vs AI Feedback (ACL Main '25)☆29Jul 23, 2025Updated last year
- ☆15Jun 17, 2025Updated last year
- Reasoning Activation in LLMs via Small Model Transfer (NeurIPS 2025)☆22Oct 16, 2025Updated 9 months ago
- Official Code for our paper: "Language Models Learn to Mislead Humans via RLHF""☆20Oct 11, 2024Updated last year
- [EMNLP 2023] Poisoning Retrieval Corpora by Injecting Adversarial Passages https://arxiv.org/abs/2310.19156☆51Dec 14, 2023Updated 2 years ago
- NeurIPS'24 - LLM Safety Landscape☆40Oct 21, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆20May 4, 2026Updated 2 months ago
- [ACL'26 Findings] Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets☆20Jun 27, 2026Updated 3 weeks ago
- ☆41May 2, 2024Updated 2 years ago
- Source code of "Task arithmetic in the tangent space: Improved editing of pre-trained models".☆113Jun 8, 2023Updated 3 years ago
- A Python library for guardrail models evaluation.☆37Oct 9, 2025Updated 9 months ago
- Reproducible, flexible LLM evaluations☆390Mar 24, 2026Updated 4 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 10 months ago
- ☆24Apr 25, 2025Updated last year
- A library for language transfer methods and algorithms.☆16Feb 6, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal☆1,014Aug 16, 2024Updated last year
- Source Code for our ICLR'26 paper☆17Feb 22, 2026Updated 5 months ago
- An official codebase for paper " CHAMPAGNE: Learning Real-world Conversation from Large-Scale Web Videos (ICCV 23)"☆52Aug 13, 2023Updated 2 years ago
- Code, data, models for the Sherlock corpus☆62Nov 11, 2022Updated 3 years ago
- ☆22Jan 13, 2025Updated last year
- ☆35Feb 17, 2026Updated 5 months ago
- Code for the paper "Pretrained Models for Multilingual Federated Learning" at NAACL 2022☆11Aug 9, 2022Updated 3 years ago
- Official github repo for SafetyBench, a comprehensive benchmark to evaluate LLMs' safety. [ACL 2024]☆296Jul 28, 2025Updated 11 months ago
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆31Aug 19, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This is the code repository for "Uncovering Safety Risks of Large Language Models through Concept Activation Vector"☆49Oct 13, 2025Updated 9 months ago
- Official repository for "DEnsity: Open-domain Dialogue Evaluation Metric using Density Estimation (ACL2023 Findings)"☆11May 23, 2023Updated 3 years ago
- ☆20Jan 7, 2026Updated 6 months ago
- S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models☆118Feb 13, 2026Updated 5 months ago
- Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic☆32Feb 18, 2026Updated 5 months ago
- DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails☆34Feb 26, 2025Updated last year
- Github repo for NeurIPS 2024 paper "Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models"☆29Dec 21, 2025Updated 7 months ago