AI safety course at the University of Tübingen (Summer Semester 2026)
☆135Aug 20, 2026Updated 2 weeks ago
Alternatives and similar repositories for tue-ai-safety-course
Users that are interested in tue-ai-safety-course are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.☆19Jun 1, 2026Updated 3 months ago
- ☆19Nov 5, 2025Updated 10 months ago
- A suite of interpretability tasks to evaluate agents using Scribe for notebook access☆18Oct 2, 2025Updated 11 months ago
- Autoresearch for LLM adversarial attacks☆241May 7, 2026Updated 4 months ago
- [ICLR 25] A novel framework for building intrinsically interpretable LLMs with human-understandable concepts to ensure safety, reliabilit…☆33Feb 5, 2026Updated 7 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [BMVC 2025] Official Implementation of the paper "PerSense: Personalized Instance Segmentation in Dense Images"☆31Dec 18, 2025Updated 8 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- A curated list of reinforcement learning in NLP. :-)☆21Oct 30, 2021Updated 4 years ago
- ☆42Jul 3, 2026Updated 2 months ago
- ☆87Feb 25, 2025Updated last year
- SC-Adagrad, SC-RMSProp and RMSProp algorithms for training deep networks proposed in☆14Oct 5, 2018Updated 7 years ago
- Provable Robustness of ReLU networks via Maximization of Linear Regions [AISTATS 2019]☆31Jul 15, 2020Updated 6 years ago
- Source code of "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" EMNLP 2025☆17Jan 12, 2026Updated 7 months ago
- Landing page for MIB: A Mechanistic Interpretability Benchmark☆26Aug 15, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆21Feb 3, 2025Updated last year
- ☆26Nov 8, 2022Updated 3 years ago
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University☆337Feb 8, 2026Updated 7 months ago
- ☆52Feb 11, 2025Updated last year
- [EMNLP'22] Textual Manifold-based Defense Against Natural Language Adversarial Examples☆11Apr 6, 2023Updated 3 years ago
- A minimal-friction secure experience that lets agents do their work. (beta)☆59Updated this week
- A claude code skill for proofreading a paper☆40Apr 23, 2026Updated 4 months ago
- ☆30Jun 19, 2023Updated 3 years ago
- [InterSpeech 2024] Official code repository of paper titled "Bird Whisperer: Leveraging Large Pre-trained Acoustic Model for Bird Call Cl…☆39Dec 11, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Aligntune : A Modular Toolkit for Post Training Alignment of LLMs☆38Updated this week
- ☆16Aug 22, 2026Updated 2 weeks ago
- ☆33Mar 18, 2026Updated 5 months ago
- Official code for "In Search of Robust Measures of Generalization" (NeurIPS 2020)☆29Dec 22, 2020Updated 5 years ago
- Repo for Paper: Discovering Interpretable Algorithms by Decompiling Transformers to RASP☆16May 25, 2026Updated 3 months ago
- Understanding and Improving Fast Adversarial Training [NeurIPS 2020]☆96Sep 23, 2021Updated 4 years ago
- ☆22Jul 22, 2026Updated last month
- [CVPR 2024] This repository includes the official implementation our paper "Revisiting Adversarial Training at Scale"☆20Apr 21, 2024Updated 2 years ago
- ☆25Mar 30, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Collection of evals for Inspect AI☆660Updated this week
- Benchmarking Open-Ended Inference Optimization by AI Agents☆43Jul 6, 2026Updated 2 months ago
- Code for paper "Robustness of Bayesian Neural Networks to Gradient-Based Attacks"☆17Feb 26, 2024Updated 2 years ago
- A corpus of short answers written by learners of English and graded with CEFR levels☆13Dec 17, 2021Updated 4 years ago
- A research workbench for developing and testing attacks against large language models, with a focus on prompt injection vulnerabilities a…☆60Jul 24, 2026Updated last month
- Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion☆11Apr 1, 2024Updated 2 years ago
- Privacy backdoors☆50Apr 28, 2024Updated 2 years ago