The jailbreak-evaluation is an easy-to-use Python package for language model jailbreak evaluation.
☆27Nov 4, 2024Updated last year
Alternatives and similar repositories for jailbreak-evaluation
Users that are interested in jailbreak-evaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Your finetuned model's back to its original safety standards faster than you can say "SafetyLock"!☆11Oct 16, 2024Updated last year
- [ACL 2025] LongSafety: Evaluating Long-Context Safety of Large Language Models☆16Jun 18, 2025Updated last year
- Adaptive Verification of Patches at the Binary Level☆15Mar 19, 2026Updated 4 months ago
- a generic decompiler testing framework that can automatically vet the decompilation correctness on the function level.☆21Sep 12, 2024Updated last year
- A toolkit for detecting and protecting against vulnerabilities in Large Language Models (LLMs).☆153Feb 4, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A list of OSINT resources and tools that may be useful when conducting investigations related to the Kingdom of Saudi Arabia☆15May 12, 2025Updated last year
- Official repository for the paper "Gradient-based Jailbreak Images for Multimodal Fusion Models" (https//arxiv.org/abs/2410.03489)☆20Oct 22, 2024Updated last year
- Data preprocessing for CCTA☆14May 29, 2025Updated last year
- Material parsers and other tools, scripts Initially developed for Grobid Superconductor☆14Feb 21, 2025Updated last year
- Math24o: 高中奥林匹克数学竞赛测评集 High School Olympiad Mathematics Chinese Benchmark☆14Mar 27, 2025Updated last year
- aigc evals☆10Dec 2, 2023Updated 2 years ago
- 原稿用紙;原稿紙;稿紙;日式便箋;UPTEX/UPLATEX 縱書☆10Nov 27, 2019Updated 6 years ago
- DefectDojo Community Content☆19Nov 9, 2025Updated 8 months ago
- lightsmile个人的用于爬取网络公开语料数据的mini通用爬虫框架。☆13Sep 30, 2020Updated 5 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- QLoRA: Efficient Finetuning of Quantized LLMs☆11Jul 22, 2023Updated 3 years ago
- The repo of the Doc2SoarGraph framework☆10Sep 17, 2024Updated last year
- Go(od) Job is a simple job scheduler that supports task retries, logging, and task sharding.☆12Jul 10, 2026Updated last week
- 🧠 LLMFuzzer - Fuzzing Framework for Large Language Models 🧠 LLMFuzzer is the first open-source fuzzing framework specifically designed …☆370Feb 12, 2024Updated 2 years ago
- A tool to create and sync web slideshows with Figma☆21Jan 4, 2025Updated last year
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated last year
- [TMLR 2025] Official implementation of AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation☆27Jun 17, 2025Updated last year
- ☆10Jun 13, 2020Updated 6 years ago
- A benchmark dataset for evaluating dialog system and natural language generation metrics.☆39Jun 13, 2022Updated 4 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [Computer Speech & Language] A transformer-based spelling error correction framework for Bangla and resource scarce Indic languages☆14Aug 9, 2024Updated last year
- Securing LLM's Against Top 10 OWASP Large Language Model Vulnerabilities 2024☆23May 10, 2024Updated 2 years ago
- Mac trackpad multi-touch to web app input piping and visualization tool☆21Oct 6, 2025Updated 9 months ago
- Sentiment Lexicon Construction☆10Sep 17, 2019Updated 6 years ago
- ☆10Feb 22, 2023Updated 3 years ago
- Docker + CVE-2015-2925 = escaping from --volume☆11Jun 30, 2015Updated 11 years ago
- [COLING 2025] Official repo of paper: "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jail…☆12Jul 26, 2024Updated last year
- NAACL 2022 paper on Analyzing Modality Robustness in Multimodal Sentiment Analysis☆31Jan 21, 2023Updated 3 years ago
- A prompt defence is a multi-layer defence that can be used to protect your applications against prompt injection attacks.☆22Apr 8, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The official implement of paper S2-VER: Semi-Supervised Visual Emotion Recognition☆11Apr 28, 2024Updated 2 years ago
- ☆12Sep 23, 2024Updated last year
- ⚡ Vigil ⚡ Detect prompt injections, jailbreaks, and other potentially risky Large Language Model (LLM) inputs☆491Jan 31, 2024Updated 2 years ago
- Supporting code for the EMNLP 2019 paper "Answers Unite! Unsupervised Metrics for Reinforced Summarization Models"☆14Jun 12, 2023Updated 3 years ago
- ☆12Jun 19, 2025Updated last year
- Finds Domain Controller on a network, enumerates users, AS-REP Roasting and hash cracking, bruteforces password, dumps AD users, DRSUAPI,…☆18Sep 23, 2023Updated 2 years ago
- ☆12Nov 3, 2022Updated 3 years ago