A verified version of the WebArena Benchmark
☆60Mar 8, 2026Updated 6 months ago
Alternatives and similar repositories for webarena-verified
Users that are interested in webarena-verified are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and re…☆638Jul 17, 2026Updated 2 months ago
- An autonomous LLM Web Agent☆16Feb 13, 2026Updated 7 months ago
- COLM2026☆38Jul 9, 2026Updated 2 months ago
- Agent Skill Induction: "Inducing Programmatic Skills for Agentic Tasks"☆47Apr 24, 2025Updated last year
- VisualWebArena is a benchmark for multimodal agents.☆487Nov 9, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?☆274Updated this week
- [ICLR 2026] Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents☆68Feb 26, 2026Updated 6 months ago
- ☆62May 26, 2026Updated 3 months ago
- Codebase for EnterpriseOps-Gym from ServiceNow☆131Aug 20, 2026Updated last month
- [EMNLP 2025] WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning☆98Nov 4, 2025Updated 10 months ago
- Scalable annotation pipeline for action-aglined fine-grained instruciton for Visual-language-Action model☆94Sep 2, 2026Updated 3 weeks ago
- ☆15May 8, 2021Updated 5 years ago
- Drive OSS standards and tools for data curation and evaluation creation for state of the art AI agents☆56Aug 11, 2026Updated last month
- ☆23Dec 21, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for paper "Prompt Engineering a Prompt Engineer" (https://arxiv.org/abs/2311.05661)☆12Aug 1, 2024Updated 2 years ago
- Code for our 2023 IEEE S&P Paper "The Leaky Web: Automated Discovery of Cross-Site Information Leaks in Browsers and the Web"☆16Jul 16, 2026Updated 2 months ago
- ☆14Dec 25, 2024Updated last year
- [ACL'25 (Findings)] Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents☆29Feb 17, 2026Updated 7 months ago
- [NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judge☆114Updated this week
- ☆12Apr 14, 2023Updated 3 years ago
- ☆17Apr 29, 2025Updated last year
- N/A☆19Aug 15, 2022Updated 4 years ago
- ☆29Apr 2, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆18Nov 1, 2024Updated last year
- Towards Large Multimodal Models as Visual Foundation Agents☆278Apr 24, 2025Updated last year
- Official Implementation of the Paper "Let's Predict Sentence by Sentence"☆15Dec 20, 2025Updated 9 months ago
- Code and dataset "ZEST" from "Learning from task descriptions", Weller et al, EMNLP 2020☆17Mar 15, 2021Updated 5 years ago
- [ICLR 2026] AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents☆52Apr 17, 2026Updated 5 months ago
- [ICLR'26 Oral] RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments☆65Feb 9, 2026Updated 7 months ago
- Scale digital agent rollouts without pain.☆36Jun 18, 2026Updated 3 months ago
- An agent framework for building and evaluating general digital agents.☆43Apr 21, 2026Updated 5 months ago
- Replication package for ESEC/FSE-2019 submission titled Diversity Web Test Generation☆15Feb 13, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [ACL 2024] On the Multi-turn Instruction Following for Conversational Web Agents☆17Oct 12, 2024Updated last year
- Challenges for general-purpose web-browsing AI agents☆69Jun 2, 2025Updated last year
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆33Apr 2, 2026Updated 5 months ago
- Awesome GUI Agent Paper List☆906Updated this week
- ☆12Jul 16, 2024Updated 2 years ago
- ☆59Apr 13, 2026Updated 5 months ago
- AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents☆61Jan 28, 2025Updated last year