[COLM 2026] Official implementation for "MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models"
☆22Sep 9, 2026Updated this week
Alternatives and similar repositories for MonitorBench
Users that are interested in MonitorBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2025] Official implementation for "The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for L…☆15May 23, 2025Updated last year
- SVIP: Towards Verifiable Inference of Open-Source Large Language Models☆15Jun 3, 2025Updated last year
- https://scale.com/research/mrt☆20Mar 16, 2026Updated 5 months ago
- Open-sourced evaluation suite from the Monitoring Monitorability paper☆98Aug 17, 2026Updated 3 weeks ago
- All-in-One Safety Evaluation Framwork☆55Aug 12, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆59Jul 29, 2026Updated last month
- [CVPR 2025] Official implementation for "Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbre…☆63Jul 5, 2025Updated last year
- Code for NeurIPS 2024 paper "Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs"☆48Feb 20, 2025Updated last year
- ☆30Jun 9, 2026Updated 3 months ago
- [ACL 2024] Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization☆29Jul 9, 2024Updated 2 years ago
- Fork of Microsoft/LightGBM to include support for the CEGB (Cost Efficient Gradient Boosting) algorithm. Original repository at https://g…☆13Jun 30, 2017Updated 9 years ago
- Long-Horizon Motion Planning with Branch-and-Bound and Neural Dynamics☆20Mar 16, 2025Updated last year
- CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR☆22Apr 7, 2026Updated 5 months ago
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆19Aug 31, 2026Updated 2 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆15Dec 7, 2021Updated 4 years ago
- [EMNLP 2025 Main] AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time☆89Jun 10, 2025Updated last year
- COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs☆18Apr 7, 2026Updated 5 months ago
- [ICLR 2025] Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron☆36Apr 30, 2025Updated last year
- [WIP] [NeurIPS 2025 Spotlight] Angular Steering: Behavior Control via Rotation in Activation Space☆25May 25, 2026Updated 3 months ago
- ☆36Jun 13, 2025Updated last year
- Code of On L-p Robustness of Decision Stumps and Trees, ICML 2020☆10Aug 3, 2020Updated 6 years ago
- Code for paper OpenWebRL: Online Multi-Turn Reinforcement Learning for Visual Web Agents☆51Aug 17, 2026Updated 3 weeks ago
- CoPur: Certifiably Robust Collaborative Inference via Feature Purification (NeurIPS 2022)☆11Dec 7, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LLM Safeguarding with Internal Representations☆20Apr 27, 2026Updated 4 months ago
- Official implementation of the KDD'26 paper "ManCAR: Manifold-Constrained Latent Reasoning with Adaptive Test-Time Computation for Sequen…☆23May 28, 2026Updated 3 months ago
- Benchmarks for the VNN Comp 2023☆16Jun 7, 2024Updated 2 years ago
- Linear and interval bound propagation in Pytorch with easy-to-use API and GPU support.☆11Aug 20, 2026Updated 3 weeks ago
- ☆11Oct 25, 2024Updated last year
- ⚛ MOSAIC is a visual-first platform that enables researchers to configure, run, and compare experiments results across RL, LLM, VLM, and …☆32Updated this week
- ☆23Jun 16, 2026Updated 2 months ago
- ☆34Mar 16, 2025Updated last year
- [ICML 2026] An Evaluation Suite for Chain-of-Thought Controllability☆53Mar 10, 2026Updated 6 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆18Feb 14, 2026Updated 7 months ago
- BitcoinEC integration/staging tree☆14Mar 22, 2017Updated 9 years ago
- [ACL 2025] LongSafety: Evaluating Long-Context Safety of Large Language Models☆16Jun 18, 2025Updated last year
- ☆11Jun 20, 2023Updated 3 years ago
- Chrome extension that logs all AJAX (XMLHttpRequest) activity to the Dev Tools Console, allowing inspection of AJAX calls, and open calls…☆26Aug 20, 2015Updated 11 years ago
- 🧨 TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?☆83Nov 27, 2025Updated 9 months ago
- The implementatioin code of paper: “A Practical Clean-Label Backdoor Attack with Limited Information in Vertical Federated Learning”☆11Jul 1, 2023Updated 3 years ago