The jailbreak-evaluation is an easy-to-use Python package for language model jailbreak evaluation.
☆28Nov 4, 2024Updated last year
Alternatives and similar repositories for jailbreak-evaluation
Users that are interested in jailbreak-evaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Your finetuned model's back to its original safety standards faster than you can say "SafetyLock"!☆11Oct 16, 2024Updated last year
- [ACL 2025] LongSafety: Evaluating Long-Context Safety of Large Language Models☆16Jun 18, 2025Updated last year
- A toolkit for detecting and protecting against vulnerabilities in Large Language Models (LLMs).☆156Feb 4, 2026Updated 7 months ago
- BAD: BiAs Detection for Large Language Models in the context of candidate screening (EECS 692)☆12Feb 14, 2024Updated 2 years ago
- TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice☆24Nov 24, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official repository for the paper "Gradient-based Jailbreak Images for Multimodal Fusion Models" (https//arxiv.org/abs/2410.03489)☆20Oct 22, 2024Updated last year
- Multi-agent system (MAS) hijacking demos☆50May 13, 2026Updated 4 months ago
- Material parsers and other tools, scripts Initially developed for Grobid Superconductor☆14Feb 21, 2025Updated last year
- ☆16Jun 20, 2022Updated 4 years ago
- aigc evals☆10Dec 2, 2023Updated 2 years ago
- [ACL 2024 main] Aligning Large Language Models with Human Preferences through Representation Engineering (https://aclanthology.org/2024.…☆28Sep 25, 2024Updated last year
- Arxiv + Notion Sync☆21May 12, 2025Updated last year
- DefectDojo Community Content☆19Aug 9, 2026Updated last month
- QLoRA: Efficient Finetuning of Quantized LLMs☆11Jul 22, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Python arbitrage bot for multiple Dex.☆11Sep 5, 2022Updated 4 years ago
- The repo of the Doc2SoarGraph framework☆11Sep 17, 2024Updated 2 years ago
- valve source engine hooking on OS X using libembryo, no sdk required☆10Sep 21, 2016Updated 10 years ago
- 🧠 LLMFuzzer - Fuzzing Framework for Large Language Models 🧠 LLMFuzzer is the first open-source fuzzing framework specifically designed …☆380Feb 12, 2024Updated 2 years ago
- Multi-head Recurrent Layer Attention for Vision Network☆23Mar 2, 2023Updated 3 years ago
- ☆17Jun 25, 2025Updated last year
- [TMLR 2025] Official implementation of AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation☆27Jun 17, 2025Updated last year
- Atlabs is a storytelling-first video creation platform that helps businesses craft engaging videos in minutes using AI. No creative or te…☆11Jul 8, 2024Updated 2 years ago
- [Computer Speech & Language] A transformer-based spelling error correction framework for Bangla and resource scarce Indic languages☆14Sep 12, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Securing LLM's Against Top 10 OWASP Large Language Model Vulnerabilities 2024☆23May 10, 2024Updated 2 years ago
- Agentic RL post-training framework☆36May 28, 2026Updated 3 months ago
- ☆21Feb 29, 2024Updated 2 years ago
- Generate security policies and documents based on KPNs templates.☆41Oct 7, 2019Updated 6 years ago
- HFT Triangular arbitrage analysis package☆12Dec 19, 2023Updated 2 years ago
- Codebase for multilingual neural machine translation☆13Nov 24, 2022Updated 3 years ago
- Sentiment Lexicon Construction☆10Sep 17, 2019Updated 7 years ago
- ☆10Feb 22, 2023Updated 3 years ago
- Transfer Learning in Dialogue Benchmarking Toolkit☆14Mar 31, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Resources for our AAAI 2022 paper: "LOREN: Logic-Regularized Reasoning for Interpretable Fact Verification".☆48Dec 14, 2022Updated 3 years ago
- NAACL 2022 paper on Analyzing Modality Robustness in Multimodal Sentiment Analysis☆31Jan 21, 2023Updated 3 years ago
- A prompt defence is a multi-layer defence that can be used to protect your applications against prompt injection attacks.☆22Apr 8, 2026Updated 5 months ago
- ☆12Sep 23, 2024Updated last year
- ☆20Oct 4, 2022Updated 3 years ago
- ⚡ Vigil ⚡ Detect prompt injections, jailbreaks, and other potentially risky Large Language Model (LLM) inputs☆496Jan 31, 2024Updated 2 years ago
- ☆25Mar 4, 2022Updated 4 years ago