A simple evaluation of generative language models and safety classifiers.
☆105Jun 16, 2026Updated last month
Alternatives and similar repositories for safety-eval
Users that are interested in safety-eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An official codebase for "NormLens: Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Comm…☆10May 9, 2024Updated 2 years ago
- Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs☆135Dec 2, 2024Updated last year
- This repository contains code for the paper "Meet Your Favorite Character: Open-domain Chatbot Mimicking Fictional Characters with only a…☆13Jun 11, 2022Updated 4 years ago
- ☆42Aug 10, 2024Updated 2 years ago
- SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types☆26Nov 29, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Corpus to accompany: "Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding"☆11Apr 11, 2025Updated last year
- Developing AGI technologies that integrate tactile and visual information to simulate human-level haptic perception and generate immersiv…☆15Dec 18, 2025Updated 7 months ago
- 【ACL 2024】 SALAD benchmark & MD-Judge☆176Mar 8, 2025Updated last year
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- Learning to route instances for Human vs AI Feedback (ACL Main '25)☆30Jul 23, 2025Updated last year
- Reasoning Activation in LLMs via Small Model Transfer (NeurIPS 2025)☆22Oct 16, 2025Updated 9 months ago
- [EMNLP 2025] The official implementation of "Zero-shot Multimodal Document Retrieval via Cross-Modal Question Generation"☆15Aug 26, 2025Updated 11 months ago
- Official Code for our paper: "Language Models Learn to Mislead Humans via RLHF""☆20Oct 11, 2024Updated last year
- [EMNLP 2023] Poisoning Retrieval Corpora by Injecting Adversarial Passages https://arxiv.org/abs/2310.19156☆51Dec 14, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- NeurIPS'24 - LLM Safety Landscape☆40Oct 21, 2025Updated 9 months ago
- ☆21May 4, 2026Updated 3 months ago
- [ACL'26 Findings] Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets☆20Jun 27, 2026Updated last month
- [NeurIPS 2025] MergeBench: A Benchmark for Merging Domain-Specialized LLMs☆49Feb 11, 2026Updated 6 months ago
- ☆15Jun 2, 2026Updated 2 months ago
- ☆41May 2, 2024Updated 2 years ago
- A Python library for guardrail models evaluation.☆39Oct 9, 2025Updated 10 months ago
- Reproducible, flexible LLM evaluations☆391Mar 24, 2026Updated 4 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆24Apr 25, 2025Updated last year
- Source Code for our ICLR'26 paper☆17Feb 22, 2026Updated 5 months ago
- An official codebase for paper " CHAMPAGNE: Learning Real-world Conversation from Large-Scale Web Videos (ICCV 23)"☆52Aug 13, 2023Updated 3 years ago
- Code, data, models for the Sherlock corpus☆62Nov 11, 2022Updated 3 years ago
- ☆22Jan 13, 2025Updated last year
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆31Aug 19, 2025Updated 11 months ago
- [ICML 2024] Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications☆91Mar 30, 2025Updated last year
- This is the code repository for "Uncovering Safety Risks of Large Language Models through Concept Activation Vector"☆49Oct 13, 2025Updated 10 months ago
- Official repository for "DEnsity: Open-domain Dialogue Evaluation Metric using Density Estimation (ACL2023 Findings)"☆11May 23, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆25Jan 7, 2026Updated 7 months ago
- S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models☆120Feb 13, 2026Updated 6 months ago
- DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails☆34Feb 26, 2025Updated last year
- Github repo for NeurIPS 2024 paper "Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models"☆29Dec 21, 2025Updated 7 months ago
- [CVPR'26] When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought☆32Feb 14, 2026Updated 6 months ago
- ☆10Mar 13, 2023Updated 3 years ago
- [ICLR 2026] Official Implementation of ProxyThinker: Test-Time Guidance through Small Visual Reasoners.☆22Sep 24, 2025Updated 10 months ago