π Benchmark your browser agent on ~2.5k READ and ACTION based tasks
β98Jul 29, 2025Updated last year
Alternatives and similar repositories for WebBench
Users that are interested in WebBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Challenges for general-purpose web-browsing AI agentsβ68Jun 2, 2025Updated last year
- Uncertainty Quantification with Pre-trained Language Models: An Empirical Analysisβ15Oct 11, 2022Updated 3 years ago
- β19Mar 7, 2026Updated 4 months ago
- β26Jul 31, 2025Updated 11 months ago
- β28Mar 12, 2022Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesisβ172Jun 18, 2026Updated last month
- The Browser Arenaβ28Updated this week
- [NeurIPS 2025]"Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning"β108Oct 21, 2025Updated 9 months ago
- ππͺ BrowserGym, a Gym environment for web task automationβ1,296Jul 17, 2026Updated last week
- RL environments + evals for AI agents. Define once, train anything.β279Updated this week
- An Illusion of Progress? Assessing the Current State of Web Agentsβ192Jun 25, 2026Updated last month
- Benchmark of complex, multimodal desktop-oriented tasks for advanced GUI-navigation AI agentsβ24May 7, 2025Updated last year
- β15Dec 21, 2025Updated 7 months ago
- Use strategy in stock transaction for high revenue.β10Dec 24, 2015Updated 10 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β33May 8, 2025Updated last year
- Companion code for FanOutQA: Multi-Hop, Multi-Document Question Answering for Large Language Models (ACL 2024)β62Apr 3, 2026Updated 3 months ago
- Public repository for "Think Twice: Perspective-Taking Improves Large Language Modelsβ Theory-of-Mind Capabilities".β25Aug 16, 2023Updated 2 years ago
- Overview of Clone Detection Tools for Javaβ14Aug 23, 2025Updated 11 months ago
- Source code for "An Empirical Study of Code Smells in Transformer-based Code Generation Techniques".β11Oct 4, 2022Updated 3 years ago
- GPT Table Semantic Parsing with complex & non-intuitive structure.β17Jul 16, 2025Updated last year
- This is for EMNLP 2024 Paper: AppBench: Planning of Multiple APIs from Various APPs for Complex User Instructionβ16Nov 4, 2024Updated last year
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.β17Updated this week
- β18Jan 3, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- β10Nov 6, 2024Updated last year
- β12Mar 24, 2023Updated 3 years ago
- β311Jul 1, 2026Updated 3 weeks ago
- A Prompt Learning Framework for Source Code Summarizationβ14Dec 26, 2023Updated 2 years ago
- GUI Grounding for Professional High-Resolution Computer Useβ383Jun 17, 2026Updated last month
- Portal: GUI Tools for Agentsβ25Sep 18, 2025Updated 10 months ago
- [ICLR'24 spotlight] Tool-Augmented Reward Modelingβ54Jun 6, 2025Updated last year
- Code for "WebVoyager: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models"β1,110Mar 4, 2024Updated 2 years ago
- Library for the Test-based Calibration Error (TCE) metric to quantify the degree to classifier calibration.β14Sep 15, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- VisualWebArena is a benchmark for multimodal agents.β484Nov 9, 2024Updated last year
- Awesome workflow-use conceptual steps & promptsβ26May 19, 2025Updated last year
- a Python library that uses Reinforcement Learning (RL) to train LLMs.β43Jul 12, 2026Updated 2 weeks ago
- Code for Findings of ACL 2021 paper "Addressing Inquiries about History: An Efficient and Practical Framework for Evaluating Open-domain β¦β19Dec 16, 2022Updated 3 years ago
- a benchmark to evaluate the situated inductive reasoningβ16Jan 7, 2025Updated last year
- β11May 23, 2023Updated 3 years ago
- VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seekingβ88Jan 21, 2026Updated 6 months ago