allenai / safety-evalLinks
A simple evaluation of generative language models and safety classifiers.
☆57Updated 11 months ago
Alternatives and similar repositories for safety-eval
Users that are interested in safety-eval are comparing it to the libraries listed below
Sorting:
- Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs☆84Updated 7 months ago
- Röttger et al. (NAACL 2024): "XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models"☆101Updated 4 months ago
- Improving Alignment and Robustness with Circuit Breakers