A simple evaluation of generative language models and safety classifiers.
☆110Aug 31, 2026Updated 3 weeks ago
Alternatives and similar repositories for safety-eval
Users that are interested in safety-eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs☆139Dec 2, 2024Updated last year
- ☆18Mar 10, 2026Updated 6 months ago
- ☆43Aug 10, 2024Updated 2 years ago
- SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types☆26Nov 29, 2024Updated last year
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆11Oct 31, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Corpus to accompany: "Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding"☆11Apr 11, 2025Updated last year
- Developing AGI technologies that integrate tactile and visual information to simulate human-level haptic perception and generate immersiv…☆15Dec 18, 2025Updated 9 months ago
- 【ACL 2024】 SALAD benchmark & MD-Judge☆175Mar 8, 2025Updated last year
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- Learning to route instances for Human vs AI Feedback (ACL Main '25)☆30Jul 23, 2025Updated last year
- ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions☆15Jun 28, 2026Updated 2 months ago
- Reasoning Activation in LLMs via Small Model Transfer (NeurIPS 2025)☆22Oct 16, 2025Updated 11 months ago
- Official Code for our paper: "Language Models Learn to Mislead Humans via RLHF""☆20Oct 11, 2024Updated last year
- NeurIPS'24 - LLM Safety Landscape☆41Oct 21, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆21May 4, 2026Updated 4 months ago
- ☆15Jun 2, 2026Updated 3 months ago
- ☆42May 2, 2024Updated 2 years ago
- A Python library for guardrail models evaluation.☆39Oct 9, 2025Updated 11 months ago
- Reproducible, flexible LLM evaluations☆395Mar 24, 2026Updated 6 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated last year
- Official code for FAccT'21 paper "Fairness Through Robustness: Investigating Robustness Disparity in Deep Learning" https://arxiv.org/abs…☆13Mar 9, 2021Updated 5 years ago
- A curated list of materials on AI guardrails☆66Jul 30, 2026Updated last month
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal☆1,053Aug 16, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Source Code for our ICLR'26 paper☆18Feb 22, 2026Updated 7 months ago
- An official codebase for paper " CHAMPAGNE: Learning Real-world Conversation from Large-Scale Web Videos (ICCV 23)"☆52Aug 13, 2023Updated 3 years ago
- Code, data, models for the Sherlock corpus☆62Nov 11, 2022Updated 3 years ago
- ☆36Feb 17, 2026Updated 7 months ago
- Code for the paper "Pretrained Models for Multilingual Federated Learning" at NAACL 2022☆11Aug 9, 2022Updated 4 years ago
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆32Aug 19, 2025Updated last year
- IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning☆19Aug 16, 2025Updated last year
- ☆45Oct 21, 2025Updated 11 months ago
- PsychAdapter is a modular framework for steering LLMs to reflect specific Big Five personality traits and mental health states using para…☆39Jun 28, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This is the code repository for "Uncovering Safety Risks of Large Language Models through Concept Activation Vector"☆49Oct 13, 2025Updated 11 months ago
- ☆30Jan 7, 2026Updated 8 months ago
- Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic☆33Feb 18, 2026Updated 7 months ago
- DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails☆34Feb 26, 2025Updated last year
- Github repo for NeurIPS 2024 paper "Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models"☆30Dec 21, 2025Updated 9 months ago
- [CVPR'26] When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought☆33Feb 14, 2026Updated 7 months ago
- ☆10Mar 13, 2023Updated 3 years ago