Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
☆556Aug 26, 2026Updated 3 weeks ago
Alternatives and similar repositories for meta-agents-research-environments
Users that are interested in meta-agents-research-environments are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Gym for Agentic LLMs☆509Jan 21, 2026Updated 7 months ago
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,328Updated this week
- The raw UserRL repo under construction☆120Jun 2, 2026Updated 3 months ago
- ☆311Jul 1, 2026Updated 2 months ago
- An interface library for RL post training with environments.☆2,596Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Our library for RL environments + evals☆4,634Updated this week
- MCP Atlas☆155Updated this week
- Research code artifacts for Code World Model (CWM) including inference tools, reproducibility, and documentation.☆908Jul 17, 2026Updated 2 months ago
- Framework for evaluating and improving agents☆5,411Updated this week
- Democratizing Reinforcement Learning for LLMs☆5,828Sep 12, 2026Updated last week
- A version of verl to support diverse tool use [TMLR 2026]☆1,041Jul 15, 2026Updated 2 months ago
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,509Apr 17, 2026Updated 5 months ago
- slime is an LLM post-training framework for RL Scaling.☆8,507Updated this week
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆74Jan 28, 2026Updated 7 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆565Sep 10, 2026Updated last week
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆487Aug 18, 2026Updated last month
- ☆1,190Jan 10, 2026Updated 8 months ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,747Apr 24, 2026Updated 4 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework