This is the repository for paper EscapeBench: Pushing Language Models to Think Outside the Box
☆18Dec 19, 2024Updated last year
Alternatives and similar repositories for EscapeBench
Users that are interested in EscapeBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- What if you need more exercises?☆36Jul 16, 2024Updated 2 years ago
- [NeurIPS'22] Trap and Replace: Defending Backdoor Attacks by Trapping Them into an Easy-to-Replace Subnetwork. Haotao Wang, Junyuan Hong,…☆15Nov 27, 2023Updated 2 years ago
- Code for paper OpenWebRL: Online Multi-Turn Reinforcement Learning for Visual Web Agents☆51Aug 17, 2026Updated last month
- Code for Evaluating Explanations for Reading Comprehension with Realistic Counterfactuals.☆17Apr 25, 2021Updated 5 years ago
- ☆24Mar 1, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This is the repository for paper "CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models"☆31Oct 8, 2023Updated 2 years ago
- AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents☆61Jan 28, 2025Updated last year
- [ICLR 2024] DMBP: Diffusion Model-Based Predictor for Robust Offline Reinforcement Learning against State Observations Perturbations.☆17May 24, 2024Updated 2 years ago
- ☆15Apr 19, 2021Updated 5 years ago
- ☆20Aug 7, 2025Updated last year
- The official repo for the code and data of paper SMART☆44Feb 20, 2025Updated last year
- ☆11Apr 29, 2019Updated 7 years ago
- ☆18Jan 3, 2022Updated 4 years ago
- ☆28Oct 18, 2022Updated 3 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆16Jul 29, 2025Updated last year
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation (ICLR 2026)☆22Apr 27, 2026Updated 4 months ago
- [ICLR 2024 Spotlight] Code for ICLR 2024 paper "Towards Robust Offline Reinforcement Learning under Diverse Data Corruption"☆22Nov 25, 2024Updated last year
- Model for processing text sequences with coreference annotations☆14Nov 29, 2018Updated 7 years ago
- Modular-HER is revised from OpenAI baselines and supports many improvements for Hindsight Experience Replay as modules.☆17Jun 23, 2021Updated 5 years ago
- ☆28May 28, 2025Updated last year
- TensorFlow implementation of the paper `Adversarial Multi-task Learning for Text Classification`☆11Apr 11, 2018Updated 8 years ago
- A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language Models☆23Jun 2, 2024Updated 2 years ago
- ☆12Nov 9, 2018Updated 7 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Text Adventure Learning Environment Suite - Benchmark to evaluate language models on interactive text environments.☆31Sep 10, 2026Updated last week
- ☆13Updated this week
- Open-source repository for the OOPSLA'24 paper "CYCLE: Learning to Self-Refine Code Generation"☆10Mar 8, 2024Updated 2 years ago
- ☆14Jul 5, 2024Updated 2 years ago
- ☆20Mar 12, 2025Updated last year
- A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models☆29Nov 25, 2024Updated last year
- The official code repository for the paper "CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments…☆35Jun 14, 2026Updated 3 months ago
- ☆13Sep 26, 2024Updated last year
- ☆35Oct 8, 2025Updated 11 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆20Oct 25, 2022Updated 3 years ago
- ☆14Jun 10, 2019Updated 7 years ago
- [NeurIPS 2024] A task generation and model evaluation system for multimodal language models.☆71Nov 27, 2024Updated last year
- code release☆44Jun 22, 2026Updated 3 months ago
- Code for ACL 2018 paper "Discourse Marker Augmented Network with Reinforcement Learning for Natural Language Inference".☆17Aug 5, 2018Updated 8 years ago
- Reproduction Code for Paper "Investigating Multi-Hop Factual Shortcuts in Knowledge Editing of Large Language Models"☆14Jun 1, 2024Updated 2 years ago
- Code for NeurIPS 2022 paper "Robust offline Reinforcement Learning via Conservative Smoothing"☆23Feb 15, 2023Updated 3 years ago