π Benchmark your browser agent on ~2.5k READ and ACTION based tasks
β99Jul 29, 2025Updated last year
Alternatives and similar repositories for WebBench
Users that are interested in WebBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Open leaderboard for browser agentsβ48Updated this week
- Challenges for general-purpose web-browsing AI agentsβ69Jun 2, 2025Updated last year
- Uncertainty Quantification with Pre-trained Language Models: An Empirical Analysisβ15Oct 11, 2022Updated 3 years ago
- This repository contains the ToolSelect dataset which was used to fine-tune Llama-2 70B for tool selection.β23Mar 11, 2024Updated 2 years ago
- β13Apr 16, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β26Jul 31, 2025Updated last year
- Opensource benchmark evaluating web operators/agents performanceβ47Apr 11, 2025Updated last year
- β28Mar 12, 2022Updated 4 years ago
- [NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesisβ179Jun 18, 2026Updated 3 months ago
- Resources for the Document Visual AI with FiftyOneβWhen a Pixel is Worth a Thousand Tokens Workshopβ19Nov 14, 2025Updated 10 months ago
- [NeurIPS 2025]"Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning"β110Oct 21, 2025Updated 11 months ago
- ππͺ BrowserGym, a Gym environment for web task automationβ1,383Updated this week
- RL environments + evals for AI agents. Define once, train anything.β304Updated this week
- OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agentsβ27May 17, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- An Illusion of Progress? Assessing the Current State of Web Agentsβ203Jun 25, 2026Updated 3 months ago
- β13Feb 22, 2024Updated 2 years ago
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applicationsβ23Oct 17, 2025Updated 11 months ago
- Benchmark of complex, multimodal desktop-oriented tasks for advanced GUI-navigation AI agentsβ24May 7, 2025Updated last year
- [LREC-Coling 2024] PECC: Problem Extraction and Coding Challengesβ14May 30, 2024Updated 2 years ago
- Tiny institutional ineptitude tracker.β17Sep 7, 2024Updated 2 years ago
- Self-hosted deep research agent with durable execution. Pi Agent SDK + Absurd + Steel.β22Jun 11, 2026Updated 3 months ago
- Use strategy in stock transaction for high revenue.β10Dec 24, 2015Updated 10 years ago
- One implementation of the paper "Controllable Neural Dialogue Summarization with Personal Named Entity Planning" (EMNLP 2022).β18Nov 9, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β18Oct 16, 2024Updated last year
- Convert BART models to ONNX with quantization. 3X reduction in size, and upto 3X boost in inference speedβ33Dec 11, 2024Updated last year
- Test how well AI agents interact with CLI toolsβ26Apr 15, 2026Updated 5 months ago
- β36Aug 26, 2026Updated last month
- [NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judgeβ114Sep 25, 2026Updated last week
- the datasets of our paperβ11Feb 26, 2024Updated 2 years ago
- Source code for "An Empirical Study of Code Smells in Transformer-based Code Generation Techniques".β11Oct 4, 2022Updated 3 years ago
- GPT Table Semantic Parsing with complex & non-intuitive structure.β17Jul 16, 2025Updated last year
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?β273Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Multi-Agent Reinforcement Learning Environment for the card game SkyJo, compatible with PettingZoo and RLLIBβ16Feb 21, 2026Updated 7 months ago
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrationsβ30May 21, 2025Updated last year
- Code for WALT βΒ Web Agents that Learn Toolsβ80Jun 2, 2026Updated 4 months ago
- β10Nov 6, 2024Updated last year
- β311Jul 1, 2026Updated 3 months ago
- A lightweight SDK for building agent-specific experiences in your app.β28May 12, 2026Updated 4 months ago
- A lightweight computational physics framework, based on the organization of turboWAVE. Implements a "Simulation, PhysicsModule, ComputeToβ¦β12Jul 23, 2026Updated 2 months ago