☆28Sep 3, 2025Updated 11 months ago
Alternatives and similar repositories for jailbreaking-frontier-models
Users that are interested in jailbreaking-frontier-models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the API, workload execution, and agents underlying the LLMail-Inject Adpative Prompt Injection Challenge☆26Apr 9, 2026Updated 4 months ago
- https://scale.com/research/mrt☆20Mar 16, 2026Updated 5 months ago
- ☆59Jul 4, 2025Updated last year
- Source code of "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" EMNLP 2025☆17Jan 12, 2026Updated 7 months ago
- ☆20Apr 10, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆10Nov 6, 2024Updated last year
- ☆49Mar 18, 2026Updated 5 months ago
- ☆127Aug 20, 2026Updated last week
- ☆17Mar 30, 2025Updated last year
- James' cookbook of evaluations and finetuning experiments☆35Feb 19, 2026Updated 6 months ago
- [ICLR 2025] Official Repository for "Tamper-Resistant Safeguards for Open-Weight LLMs"☆70Jun 9, 2025Updated last year
- ☆23May 31, 2026Updated 2 months ago
- Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial Examples☆19May 23, 2022Updated 4 years ago
- Sharpness-Aware Minimization Leads to Low-Rank Features [NeurIPS 2023]☆29Sep 22, 2023Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆30Oct 22, 2024Updated last year
- General research for Dreadnode☆28Jun 17, 2024Updated 2 years ago
- Applying SAEs for fine-grained control☆27Dec 15, 2024Updated last year
- Improving Alignment and Robustness with Circuit Breakers☆268Sep 24, 2024Updated last year
- ☆11Dec 8, 2022Updated 3 years ago
- ☆14Apr 27, 2022Updated 4 years ago
- ASSIST: Towards Label Noise-Robust Dialogue State Tracking☆10Apr 11, 2022Updated 4 years ago
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆34Jan 23, 2025Updated last year
- Saving Dense Retriever from Shortcut Dependency in Conversational Search (EMNLP 2022)☆18Nov 24, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Optimization algorithm which fits a ResNet to CIFAR-10 5x faster than SGD / Adam (with terrible generalization)☆14Oct 20, 2023Updated 2 years ago
- Neural Machine Translation model for Capstone Project☆11Apr 11, 2020Updated 6 years ago
- ☆23Jul 26, 2025Updated last year
- Official release of code for the paper RL is a hammer and LLMs are nails A simple RL approach to stronger prompt injection attacks☆53May 6, 2026Updated 3 months ago
- MishformerLens intends to be a drop-in replacement for TransformerLens that AST patches HuggingFace Transformers rather than implementing…☆10Oct 7, 2024Updated last year
- Easily deploy my zsh and tmux configuration on new machines. Includes local and remote aliases to improve workflow.☆16Apr 23, 2026Updated 4 months ago
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.☆81Jul 20, 2026Updated last month
- playing with gpt4☆13Mar 17, 2023Updated 3 years ago
- Digital texts in Prakrit☆11Sep 14, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆25Dec 18, 2025Updated 8 months ago
- ☆19Sep 20, 2022Updated 3 years ago
- On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them [NeurIPS 2020]☆36Jul 3, 2021Updated 5 years ago
- [NAACL 2025] ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage☆16Sep 2, 2025Updated 11 months ago
- Code snippets to reproduce MCP tool poisoning attacks.☆206Apr 10, 2025Updated last year
- A Blackjack game with GUI written in Java.☆11Nov 21, 2018Updated 7 years ago
- Implementation of path patching & activation patching (will eventually add to TransformerLens).☆15Jan 8, 2024Updated 2 years ago