Code for safety test in "Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates"
☆22Sep 21, 2025Updated 9 months ago
Alternatives and similar repositories for PTST
Users that are interested in PTST are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Our research proposes a novel MoGU framework that improves LLMs' safety while preserving their usability.☆18Jan 14, 2025Updated last year
- ☆14Aug 9, 2023Updated 2 years ago
- Code for SafeMERGE (ICLR 2025).☆15Apr 1, 2025Updated last year
- [EMNLP 2025] Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment☆15Jul 22, 2025Updated 11 months ago
- Code for LLM_Catastrophic_Forgetting via SAM.☆11Jun 7, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning [ICML 2024]☆21May 2, 2024Updated 2 years ago
- We jailbreak GPT-3.5 Turbo’s safety guardrails by fine-tuning it on only 10 adversarially designed examples, at a cost of less than $0.20…☆358Feb 23, 2024Updated 2 years ago
- ☆21Apr 16, 2025Updated last year
- ☆61Mar 9, 2023Updated 3 years ago
- Fine-tuning-free Shapley value (FreeShap) for instance attribution☆14May 29, 2024Updated 2 years ago
- This is the repository that introduces research topics related to protecting intellectual property (IP) of AI from a data-centric perspec…☆23Oct 30, 2023Updated 2 years ago
- Unofficial implementation of Chain of Hindsight (https://arxiv.org/abs/2302.02676) using pytorch and huggingface Trainers.☆11Apr 5, 2023Updated 3 years ago
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆33Jan 23, 2025Updated last year
- Improving Alignment and Robustness with Circuit Breakers☆266Sep 24, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Simple and scalable tools for data-driven pretraining data selection.☆30Jun 9, 2025Updated last year
- Github repo for NeurIPS 2024 paper "Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models"☆29Dec 21, 2025Updated 7 months ago
- The official implement of paper "Does Federated Learning Really Need Backpropagation?"☆23Feb 9, 2023Updated 3 years ago
- An implementation of SEAL: Safety-Enhanced Aligned LLM fine-tuning via bilevel data selection.☆24Feb 20, 2025Updated last year
- Simple template for quick prototyping and standardization of deep learning projects☆11Dec 29, 2023Updated 2 years ago
- ☆11Jul 31, 2020Updated 5 years ago
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models (ICLR 2024)☆82Oct 3, 2024Updated last year
- ☆23May 14, 2025Updated last year
- "In-Context Unlearning: Language Models as Few Shot Unlearners". Martin Pawelczyk, Seth Neel* and Himabindu Lakkaraju*; ICML 2024.☆30Oct 18, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Codebase for "On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback". This repo implements a generative multi-tur…☆25Dec 3, 2024Updated last year
- [ACL 2024] The official codebase for the paper "Self-Distillation Bridges Distribution Gap in Language Model Fine-tuning".☆167Nov 2, 2024Updated last year
- ☆35Sep 13, 2023Updated 2 years ago
- Official repository for "Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks"☆62Aug 8, 2024Updated last year
- Multi-Layer Sparse Autoencoders (ICLR 2025)☆30Feb 6, 2026Updated 5 months ago
- [ICLR 2025] On Evluating the Durability of Safegurads for Open-Weight LLMs☆13Jun 20, 2025Updated last year
- Code for paper "Universal Jailbreak Backdoors from Poisoned Human Feedback"☆67Apr 24, 2024Updated 2 years ago
- ☆33Feb 11, 2025Updated last year
- Identification of the Adversary from a Single Adversarial Example (ICML 2023)☆10Jul 15, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [AAAI26] Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilitie…☆11Feb 7, 2026Updated 5 months ago
- [ICLR 2022] Denoising Likelihood Score Matching for Conditional Score-based Data Generation☆11Jun 15, 2026Updated last month
- Code for paper "Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion"☆14Mar 28, 2024Updated 2 years ago
- A collection of judges for evaluating LLM model output for safety & toxicity with a standardized API.☆15Jan 7, 2026Updated 6 months ago
- Implementation of Evolutionary and Metropolis Hastings Monte Carlo for text based (e.g. nucleotide/peptide) sequences☆13Mar 7, 2024Updated 2 years ago
- [NeurIPS'24] "NeuralFuse: Learning to Recover the Accuracy of Access-Limited Neural Network Inference in Low-Voltage Regimes" by Hao-Lun …☆10Sep 18, 2025Updated 10 months ago
- Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections☆21Jul 2, 2026Updated 2 weeks ago