π AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
β502Sep 4, 2026Updated this week
Alternatives and similar repositories for appworld
Users that are interested in appworld are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code and Data for Tau-Benchβ1,425Mar 18, 2026Updated 5 months ago
- A new tool learning benchmark aiming at well-balanced stability and reality, based on ToolBench.β240Apr 15, 2025Updated last year
- [NeurIPS 2022] πWebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agentsβ589Sep 6, 2024Updated 2 years ago
- Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"β1,605Nov 26, 2025Updated 9 months ago
- verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-inβ¦β2,282Jun 9, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Agent RL framework for LLM agents: multi-turn reinforcement learning with StarPO and reasoning-collapse diagnosticsβ2,794Aug 23, 2026Updated 2 weeks ago
- [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environmentsβ3,129Aug 30, 2026Updated last week
- Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhiβ¦β842May 30, 2026Updated 3 months ago
- Ο-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domainsβ1,969Updated this week
- VisualWebArena is a benchmark for multimodal agents.β485Nov 9, 2024Updated last year
- [ICML'24 Spotlight] "TravelPlanner: A Benchmark for Real-World Planning with Language Agents"β544May 24, 2026Updated 3 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Frameworkβ23,336Updated this week
- AndroidWorld is an environment and benchmark for autonomous agentsβ874Jul 16, 2026Updated last month
- WebGym: Web-browser-based tasks for RL Agentsβ24Feb 4, 2021Updated 5 years ago
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- β311Jul 1, 2026Updated 2 months ago
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Executionβ479Aug 18, 2026Updated 2 weeks ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoningβ405Mar 30, 2026Updated 5 months ago
- An Analytical Evaluation Board of Multi-turn LLM Agents [NeurIPS 2024 Oral]β444May 20, 2024Updated 2 years ago
- β230Jun 2, 2025Updated last year
- A virtual environment for developing and evaluating automated scientific discovery agents.β219Mar 10, 2025Updated last year
- [NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2β152Apr 20, 2026Updated 4 months ago
- β280Nov 7, 2025Updated 10 months ago
- Official implementation of paper "ACON: Optimizing Context Compression for Long-horizon LLM Agents"β110Oct 14, 2025Updated 10 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]β732Jul 29, 2025Updated last year
- Building Open LLM Web Agents with Self-Evolving Online Curriculum RLβ539Jun 6, 2025Updated last year
- Official repo for paper DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning.β396Feb 22, 2025Updated last year
- SkyRL: A Modular Full-stack RL Library for LLMsβ2,259Updated this week
- Towards Large Multimodal Models as Visual Foundation Agentsβ276Apr 24, 2025Updated last year
- An agent benchmark with tasks in a simulated software company.β775Nov 17, 2025Updated 9 months ago
- Awesome GUI Agent Paper Listβ896Aug 17, 2026Updated 3 weeks ago
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learningβ853Feb 8, 2026Updated 7 months ago
- SWE-bench: Can Language Models Resolve Real-world Github Issues?β5,796Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICCV 2025] GUIOdyssey is a comprehensive dataset for training and evaluating cross-app navigation agents. GUIOdyssey consists of 8,834 eβ¦β161Jan 3, 2026Updated 8 months ago
- This is the repository for the Tool Learning survey.β488Aug 9, 2025Updated last year
- Reinforcement learning for improving embodied task planning in large language modelsβ28Mar 30, 2026Updated 5 months ago
- A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)β3,716Feb 8, 2026Updated 6 months ago
- [COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agentsβ331Jul 13, 2025Updated last year
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Reiβ¦β1,433May 16, 2025Updated last year
- Sotopia: an Open-ended Social Learning Environment (ICLR 2024 spotlight)β330Jun 5, 2026Updated 3 months ago