Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
☆547Aug 26, 2026Updated this week
Alternatives and similar repositories for meta-agents-research-environments
Users that are interested in meta-agents-research-environments are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Gym for Agentic LLMs☆505Jan 21, 2026Updated 7 months ago
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,210Updated this week
- The raw UserRL repo under construction☆118Jun 2, 2026Updated 2 months ago
- ☆311Jul 1, 2026Updated last month
- An interface library for RL post training with environments.☆2,526Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Our library for RL environments + evals☆4,570Updated this week
- Research code artifacts for Code World Model (CWM) including inference tools, reproducibility, and documentation.☆894Jul 17, 2026Updated last month
- MCP Atlas☆150Aug 12, 2026Updated 2 weeks ago
- Framework for evaluating and improving agents☆4,782Updated this week
- Democratizing Reinforcement Learning for LLMs☆5,808Aug 24, 2026Updated last week
- A version of verl to support diverse tool use [TMLR 2026]☆1,036Jul 15, 2026Updated last month
- slime is an LLM post-training framework for RL Scaling.☆8,313Updated this week
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,497Apr 17, 2026Updated 4 months ago
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆73Jan 28, 2026Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆538Aug 21, 2026Updated last week
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆470Aug 18, 2026Updated last week
- ☆1,186Jan 10, 2026Updated 7 months ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,725Apr 24, 2026Updated 4 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,204Updated this week
- Agent RL framework for LLM agents: multi-turn reinforcement learning with StarPO and reasoning-collapse diagnostics☆2,781Aug 23, 2026Updated last week
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆2,288Updated this week
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆145May 6, 2026Updated 3 months ago
- Harness for running and evaluating AI agents against RL environments☆248Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.☆5,704Updated this week
- Agentic RL Training at Scale☆1,993Updated this week
- AIRA-dojo: a framework for developing and evaluating AI research agents☆162Apr 14, 2026Updated 4 months ago
- MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.☆596Jun 23, 2026Updated 2 months ago
- Search, understand, reproduce, and improve an idea with ease☆1,227Updated this week
- An agent benchmark with tasks in a simulated software company.☆771Nov 17, 2025Updated 9 months ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,558Jul 11, 2026Updated last month
- Post-training with Tinker☆4,067Updated this week
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆728Jul 29, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhi…☆837May 30, 2026Updated 3 months ago
- AllenAI's post-training codebase☆3,854Updated this week
- Ideas for projects related to Tinker☆196Nov 6, 2025Updated 9 months ago
- τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains☆1,911Updated this week
- 🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.☆670Jan 29, 2026Updated 7 months ago
- MLGym A New Framework and Benchmark for Advancing AI Research Agents☆619Aug 10, 2025Updated last year
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning☆204Mar 27, 2026Updated 5 months ago