Github repo for NeurIPS 2024 paper "Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models"
☆30Dec 21, 2025Updated 9 months ago
Alternatives and similar repositories for SafeLoRA
Users that are interested in SafeLoRA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for “SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation(ICLR 2025)”☆29Oct 23, 2025Updated 11 months ago
- Code for SafeMERGE (ACL 2026).☆15Oct 2, 2026Updated last week
- [EMNLP 2025] Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment☆17Jul 22, 2025Updated last year
- ☆24Dec 8, 2024Updated last year
- Official Repository for The Paper: Safety Alignment Should Be Made More Than Just a Few Tokens Deep☆192Apr 23, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is the official code for the paper "Vaccine: Perturbation-aware Alignment for Large Language Models" (NeurIPS2024)☆53Jan 15, 2026Updated 8 months ago
- ☆20Jun 21, 2025Updated last year
- This is the official code for the paper "Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning" (NeurIPS2024)☆28Sep 10, 2024Updated 2 years ago
- NeurIPS'24 - LLM Safety Landscape☆41Oct 21, 2025Updated 11 months ago
- This is the official code for the paper "Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturba…☆43Mar 22, 2025Updated last year
- Inference-time alignment for harmlessness through cross-model guidance (ACL 2024). Code + MM-Harmful Bench.☆38Oct 2, 2024Updated 2 years ago
- ☆45Jul 3, 2026Updated 3 months ago
- Our research proposes a novel MoGU framework that improves LLMs' safety while preserving their usability.☆18Jan 14, 2025Updated last year
- ☆11Jun 20, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An implementation of SEAL: Safety-Enhanced Aligned LLM fine-tuning via bilevel data selection.☆24Feb 20, 2025Updated last year
- The first toolkit for MLRM safety evaluation, providing unified interface for mainstream models, datasets, and jailbreaking methods!☆15Apr 8, 2025Updated last year
- code space of paper "Safety Layers in Aligned Large Language Models: The Key to LLM Security" (ICLR 2025)☆26Apr 26, 2025Updated last year
- [NDSS'25] The official implementation of safety misalignment.☆19Jan 8, 2025Updated last year
- Code and dataset for the paper: "Can Editing LLMs Inject Harm?" [AAAI'26]☆21Dec 26, 2025Updated 9 months ago
- ☆49Oct 1, 2024Updated 2 years ago
- This is the official code for the paper "Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation"☆56Feb 2, 2025Updated last year
- Applies ROME and MEMIT on Mamba-S4 models☆16Apr 5, 2024Updated 2 years ago
- This is the code repository for "Uncovering Safety Risks of Large Language Models through Concept Activation Vector"☆49Oct 13, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Precision Knowledge Editing (PKE): A novel method to reduce toxicity in LLMs while preserving performance, with robust evaluations and ha…☆12Nov 26, 2024Updated last year
- ☆23May 23, 2025Updated last year
- [CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment☆29Jun 11, 2025Updated last year
- [ICLR 2025] Adaptive prompt tailored pruning of T2I diffusion models.☆15Feb 1, 2025Updated last year
- The official repository for paper "MLLM-Protector: Ensuring MLLM’s Safety without Hurting Performance"☆46Apr 21, 2024Updated 2 years ago
- Official codebase for "STAIR: Improving Safety Alignment with Introspective Reasoning"☆90Feb 26, 2025Updated last year
- The repo for HiRA paper☆37Jan 9, 2026Updated 9 months ago
- ☆35Sep 12, 2026Updated 3 weeks ago
- ☆48Jan 15, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Restore safety in fine-tuned language models through task arithmetic☆33Mar 28, 2024Updated 2 years ago
- Prompt Generator model for Stable Diffusion Models☆12Jun 20, 2023Updated 3 years ago
- [CVPR 2024] SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos☆18May 21, 2024Updated 2 years ago
- [COLM 2025] SEAL: Steerable Reasoning Calibration of Large Language Models for Free☆66Apr 6, 2025Updated last year
- PHASE annotations for societal bias in vision-and-language tasks.☆18Jun 18, 2024Updated 2 years ago
- (ACL '25 - Oral) FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Foundation Models☆37Oct 4, 2025Updated last year
- ICLR2024 Paper. Showing properties of safety tuning and exaggerated safety.☆96May 9, 2024Updated 2 years ago