AI safety course at the University of Tübingen (Summer Semester 2026)
☆130Aug 10, 2026Updated last week
Alternatives and similar repositories for tue-ai-safety-course
Users that are interested in tue-ai-safety-course are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.☆19Jun 1, 2026Updated 2 months ago
- ☆18Nov 5, 2025Updated 9 months ago
- A suite of interpretability tasks to evaluate agents using Scribe for notebook access☆18Oct 2, 2025Updated 10 months ago
- Autoresearch for LLM adversarial attacks☆240May 7, 2026Updated 3 months ago
- [BMVC 2025] Official Implementation of the paper "PerSense: Personalized Instance Segmentation in Dense Images"☆31Dec 18, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- XoRL☆33Updated this week
- ☆42Jul 3, 2026Updated last month
- ☆86Feb 25, 2025Updated last year
- Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections☆22Jul 2, 2026Updated last month
- Implementation of average- and worst-case robust flatness measures for adversarial training.☆15Nov 5, 2021Updated 4 years ago
- Provable Robustness of ReLU networks via Maximization of Linear Regions [AISTATS 2019]☆31Jul 15, 2020Updated 6 years ago
- Landing page for MIB: A Mechanistic Interpretability Benchmark☆26Aug 15, 2025Updated last year
- A toolkit for neural language modeling using Tensorflow including basic models like RNNs and LSTMs as well as more advanced models.☆21Jan 31, 2019Updated 7 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆26Nov 8, 2022Updated 3 years ago
- Course Materials for Interpretability of Large Language Models (0368.4264) at Tel Aviv University☆337Feb 8, 2026Updated 6 months ago
- ☆46Feb 11, 2025Updated last year
- [EMNLP'22] Textual Manifold-based Defense Against Natural Language Adversarial Examples☆11Apr 6, 2023Updated 3 years ago
- You've probably been using agents irresponsibly. Today, you can turn a new leaf.☆54Updated this week
- ☆25May 25, 2024Updated 2 years ago
- ☆30Jun 19, 2023Updated 3 years ago
- [InterSpeech 2024] Official code repository of paper titled "Bird Whisperer: Leveraging Large Pre-trained Acoustic Model for Bird Call Cl…☆39Dec 11, 2024Updated last year
- Aligntune : A Modular Toolkit for Post Training Alignment of LLMs☆38Jul 20, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15Dec 4, 2025Updated 8 months ago
- ☆34Mar 18, 2026Updated 5 months ago
- Repo for Paper: Discovering Interpretable Algorithms by Decompiling Transformers to RASP☆16May 25, 2026Updated 2 months ago
- Understanding and Improving Fast Adversarial Training [NeurIPS 2020]☆96Sep 23, 2021Updated 4 years ago
- ☆22Jul 22, 2026Updated 3 weeks ago
- Code for "RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models"☆17Apr 21, 2026Updated 3 months ago
- [ECCV 2024] Towards Reliable Evaluation and Fast Training of Robust Semantic Segmentation Models☆21Jul 17, 2024Updated 2 years ago
- [CVPR 2024] This repository includes the official implementation our paper "Revisiting Adversarial Training at Scale"☆20Apr 21, 2024Updated 2 years ago
- SGD with large step sizes learns sparse features [ICML 2023]☆34Apr 24, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A corpus of short answers written by learners of English and graded with CEFR levels☆12Dec 17, 2021Updated 4 years ago
- A research workbench for developing and testing attacks against large language models, with a focus on prompt injection vulnerabilities a…☆61Jul 24, 2026Updated 3 weeks ago
- Privacy backdoors☆50Apr 28, 2024Updated 2 years ago
- Code repo for the model organisms and convergent directions of EM papers.☆79Sep 22, 2025Updated 10 months ago
- ☆30Jul 1, 2026Updated last month
- Improving Alignment and Robustness with Circuit Breakers☆267Sep 24, 2024Updated last year
- Training vision models with full-batch gradient descent and regularization☆40Feb 14, 2023Updated 3 years ago