[ICLR 2025] Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron
☆36Apr 30, 2025Updated last year
Alternatives and similar repositories for Safety-Neuron
Users that are interested in Safety-Neuron are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] The implementation of paper "The Emergence of Abstract Thought in Large Language Models Beyond Any Language"☆20Jun 9, 2025Updated last year
- Code for the paper "AsFT: Anchoring Safety During LLM Fune-Tuning Within Narrow Safety Basin".☆38Jul 10, 2025Updated last year
- Confidence Regulation Neurons in Language Models (NeurIPS 2024)☆16Feb 1, 2025Updated last year
- ☆44Jul 3, 2026Updated 2 months ago
- [NeurIPS 2024] How do Large Language Models Handle Multilingualism?☆52Nov 8, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆36Jun 13, 2025Updated last year
- ☆15Aug 22, 2026Updated last month
- The official repository of 'Unnatural Language Are Not Bugs but Features for LLMs'☆25May 20, 2025Updated last year
- [ICLR 2026 Oral] Invisible Safety Threat: Malicious Finetuning for LLM via Steganography☆21Mar 22, 2026Updated 6 months ago
- [ACL 2025] Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints☆19May 23, 2025Updated last year
- [COLM 2026] Official implementation for "MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo…☆22Sep 9, 2026Updated 2 weeks ago
- ☆79Mar 6, 2025Updated last year
- Official code for the paper "Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?" (ICLR 2024)☆11Aug 26, 2024Updated 2 years ago
- Benchmarking Dark Patterns in LLMs (ICLR 2025)☆19Mar 29, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Design for Error Detection in Deep-Research Agents Trajectories.☆24Jun 4, 2026Updated 3 months ago
- [CVPR '23 Highlight] Official repository for the paper "Quantum Multi-Model Fitting".☆12Mar 7, 2025Updated last year
- [ICLR 2025] Adaptive prompt tailored pruning of T2I diffusion models.☆15Feb 1, 2025Updated last year
- [ICLR 2025] Official implementation of paper "Dynamic Low-Rank Sparse Adaptation for Large Language Models".☆25Mar 16, 2025Updated last year
- [ICML 2023] "NeRFool: Uncovering the Vulnerability of Generalizable Neural Radiance Fields against Adversarial Perturbations" by Yonggan …☆19Mar 10, 2024Updated 2 years ago
- [ICML 2024] Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications☆91Mar 30, 2025Updated last year
- ☆84Mar 30, 2025Updated last year
- ☆71Jun 1, 2025Updated last year
- ☆11Apr 2, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆48May 29, 2026Updated 3 months ago
- ☆12Sep 28, 2023Updated 3 years ago
- [ICML 2026] Official implementation for paper "Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Ag…☆38Jul 31, 2026Updated last month
- Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".☆448Jun 13, 2025Updated last year
- ☆12Jun 5, 2024Updated 2 years ago
- This repo is the official implementation of “Are Your Agents Upward Deceivers?”. The paper is accepted by ICML 2026.☆24Dec 15, 2025Updated 9 months ago
- ECSO (Make MLLM safe without neither training nor any external models!) (https://arxiv.org/abs/2403.09572)☆38Nov 2, 2024Updated last year
- [MICCAI'25] ClipGS: Clippable Gaussian Splatting for Interactive Cinematic Visualization of Volumetric Medical Data☆16Jul 28, 2025Updated last year
- K-Means algorithm in the Poincare Disk Model☆16Nov 12, 2018Updated 7 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Using PCA, Autoencoder and Fisher linear discriminant to extract the effective representations from the face images. Do the reconstructio…☆12Apr 23, 2019Updated 7 years ago
- ☆19Nov 5, 2025Updated 10 months ago
- Accepted by ECCV 2024☆222Oct 15, 2024Updated last year
- PHASE annotations for societal bias in vision-and-language tasks.☆18Jun 18, 2024Updated 2 years ago
- Graph-based experience memory for LLM reward prediction with limited labels. 20% labels → 97.3% Oracle.☆19Mar 24, 2026Updated 6 months ago
- Source code of BI-Mamba for cardiovascular disease detection from two-view chest X-rays☆16Dec 10, 2025Updated 9 months ago
- [EMNLP 2025] Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking☆12Aug 22, 2025Updated last year