ππͺ BrowserGym, a Gym environment for web task automation
β1,284Jul 17, 2026Updated this week
Alternatives and similar repositories for BrowserGym
Users that are interested in BrowserGym are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reβ¦β606Updated this week
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?β258Apr 25, 2026Updated 2 months ago
- Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"β1,550Nov 26, 2025Updated 7 months ago
- Standardize benchmark wrapping so the community can wrap various otherwise-incompatible benchmarks uniformly and use them everywhere.β52Updated this week
- VisualWebArena is a benchmark for multimodal agents.β482Nov 9, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environmentsβ3,026Updated this week
- TapeAgents is a framework that facilitates all stages of the LLM Agent development lifecycleβ318Dec 16, 2025Updated 7 months ago
- Drive OSS standards and tools for data curation and evaluation creation for state of the art AI agentsβ54Updated this week
- An Illusion of Progress? Assessing the Current State of Web Agentsβ191Jun 25, 2026Updated 3 weeks ago
- A verified version of the WebArena Benchmarkβ44Mar 8, 2026Updated 4 months ago
- A scalable asynchronous reinforcement learning implementation with in-flight weight updates.β427Updated this week
- Building Open LLM Web Agents with Self-Evolving Online Curriculum RLβ535Jun 6, 2025Updated last year
- An agent benchmark with tasks in a simulated software company.β747Nov 17, 2025Updated 8 months ago
- WebLINX is a benchmark for building web navigation agents with conversational capabilitiesβ162Feb 11, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Awesome GUI Agent Paper Listβ861Jun 28, 2026Updated 3 weeks ago
- [NeurIPS'23 Spotlight] "Mind2Web: Towards a Generalist Agent for the Web" -- the first LLM-based web agent and benchmark for generalist wβ¦β1,015Nov 5, 2025Updated 8 months ago
- AWM: Agent Workflow Memoryβ447Dec 22, 2025Updated 6 months ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]β708Jul 29, 2025Updated 11 months ago
- Code for "WebVoyager: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models"β1,110Mar 4, 2024Updated 2 years ago
- β41Jul 21, 2024Updated last year
- Code for the paper π³ Tree Search for Language Model Agentsβ223Jul 25, 2024Updated last year
- DoomArena is a Framework for Testing AI Agents Against Evolving Security Threatsβ62Sep 12, 2025Updated 10 months ago
- Setup scripts for the WebArena benchmarkβ22Jun 19, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- OS-ATLAS: A Foundation Action Model For Generalist GUI Agentsβ452Apr 20, 2025Updated last year
- [ICML'24] SeeAct is a system for generalist web agents that autonomously carry out tasks on any given website, with a focus on large multβ¦β851Feb 3, 2025Updated last year
- SkyRL: A Modular Full-stack RL Library for LLMsβ2,081Updated this week
- [NeurIPS 2022] πWebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agentsβ571Sep 6, 2024Updated last year
- COLM2026β36Jul 9, 2026Updated last week
- [NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judgeβ111May 17, 2026Updated 2 months ago
- Democratizing Reinforcement Learning for LLMsβ5,708Updated this week
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Frameworkβ22,571Updated this week
- A Benchmark for Evaluating Safety and Trustworthiness in Web Agents for Enterprise Scenariosβ25Mar 12, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Windows Agent Arena (WAA) πͺ is a scalable OS platform for testing and benchmarking of multi-modal AI agents.β881Apr 13, 2026Updated 3 months ago
- π» A curated list of papers and resources for multi-modal Graphical User Interface (GUI) agents.β1,197Aug 17, 2025Updated 11 months ago
- [ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agentsβ314Mar 11, 2026Updated 4 months ago
- Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhiβ¦β813May 30, 2026Updated last month
- [ACL'25 (Findings)] Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agentsβ29Feb 17, 2026Updated 5 months ago
- [NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesisβ172Jun 18, 2026Updated last month
- A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)β3,586Feb 8, 2026Updated 5 months ago