π Benchmark your browser agent on ~2.5k READ and ACTION based tasks
β99Jul 29, 2025Updated last year
Alternatives and similar repositories for WebBench
Users that are interested in WebBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Open leaderboard for browser agentsβ48Sep 5, 2026Updated last week
- Challenges for general-purpose web-browsing AI agentsβ69Jun 2, 2025Updated last year
- Uncertainty Quantification with Pre-trained Language Models: An Empirical Analysisβ15Oct 11, 2022Updated 3 years ago
- This repository contains the ToolSelect dataset which was used to fine-tune Llama-2 70B for tool selection.β23Mar 11, 2024Updated 2 years ago
- β13Apr 16, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β18Jan 3, 2024Updated 2 years ago
- β26Jul 31, 2025Updated last year
- [NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesisβ178Jun 18, 2026Updated 2 months ago
- A collection of particularly difficult test scenarios for evaluating browser-use.β28May 15, 2026Updated 3 months ago
- [NeurIPS 2025]"Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning"β109Oct 21, 2025Updated 10 months ago
- ππͺ BrowserGym, a Gym environment for web task automationβ1,355Jul 17, 2026Updated last month
- RL environments + evals for AI agents. Define once, train anything.β300Updated this week
- OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agentsβ27May 17, 2026Updated 3 months ago
- An Illusion of Progress? Assessing the Current State of Web Agentsβ200Jun 25, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Benchmark of complex, multimodal desktop-oriented tasks for advanced GUI-navigation AI agentsβ24May 7, 2025Updated last year
- [LREC-Coling 2024] PECC: Problem Extraction and Coding Challengesβ14May 30, 2024Updated 2 years ago
- β15Dec 21, 2025Updated 8 months ago
- Use strategy in stock transaction for high revenue.β10Dec 24, 2015Updated 10 years ago
- One implementation of the paper "Controllable Neural Dialogue Summarization with Personal Named Entity Planning" (EMNLP 2022).β18Nov 9, 2023Updated 2 years ago
- Convert BART models to ONNX with quantization. 3X reduction in size, and upto 3X boost in inference speedβ33Dec 11, 2024Updated last year
- Interface for interacting with Gradient AI in Pythonβ15Jun 28, 2024Updated 2 years ago
- Test how well AI agents interact with CLI toolsβ26Apr 15, 2026Updated 4 months ago
- WIP: Ofen is a toolkit aimed at making transformer models production-ready. API includedβ17Oct 2, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β14Jul 10, 2025Updated last year
- Overview of Clone Detection Tools for Javaβ14Aug 23, 2025Updated last year
- β17Oct 30, 2023Updated 2 years ago
- [NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judgeβ114May 17, 2026Updated 3 months ago
- Source code for "An Empirical Study of Code Smells in Transformer-based Code Generation Techniques".β11Oct 4, 2022Updated 3 years ago
- the datasets of our paperβ11Feb 26, 2024Updated 2 years ago
- This is for EMNLP 2024 Paper: AppBench: Planning of Multiple APIs from Various APPs for Complex User Instructionβ16Nov 4, 2024Updated last year
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?β271Apr 25, 2026Updated 4 months ago
- Multi-Agent Reinforcement Learning Environment for the card game SkyJo, compatible with PettingZoo and RLLIBβ16Feb 21, 2026Updated 6 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrationsβ30May 21, 2025Updated last year
- This repository is designed for deploying and managing server processes that handle embeddings using the Infinity Embedding model or Largβ¦β25Mar 6, 2025Updated last year
- β18Jan 3, 2025Updated last year
- π₯ A list of tools, frameworks, and resources for building AI web agentsβ1,570Aug 25, 2026Updated 2 weeks ago
- β10Nov 6, 2024Updated last year
- β12Mar 24, 2023Updated 3 years ago
- β311Jul 1, 2026Updated 2 months ago