Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
☆528Jun 20, 2026Updated last month
Alternatives and similar repositories for meta-agents-research-environments
Users that are interested in meta-agents-research-environments are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Gym for Agentic LLMs☆502Jan 21, 2026Updated 6 months ago
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,085Updated this week
- The raw UserRL repo under construction☆111Jun 2, 2026Updated last month
- ☆308Jul 1, 2026Updated 2 weeks ago
- An interface library for RL post training with environments.☆2,439Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Our library for RL environments + evals☆4,390Updated this week
- MCP Atlas☆121Updated this week
- Research code artifacts for Code World Model (CWM) including inference tools, reproducibility, and documentation.☆881Updated this week
- Framework for evaluating and improving agents☆3,348Updated this week
- Democratizing Reinforcement Learning for LLMs☆5,708Updated this week
- A version of verl to support diverse tool use [TMLR 2026]☆1,021Updated this week
- slime is an LLM post-training framework for RL Scaling.☆7,569Updated this week
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,463Apr 17, 2026Updated 3 months ago
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆71Jan 28, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆463Updated this week
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆432Updated this week
- ☆1,170Jan 10, 2026Updated 6 months ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,647Apr 24, 2026Updated 2 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆22,587Updated this week
- RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.☆2,753Apr 14, 2026Updated 3 months ago
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆1,761Updated this week
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆145May 6, 2026Updated 2 months ago
- Harness for running and evaluating AI agents against RL environments☆217Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Agentic RL Training at Scale☆1,702Updated this week
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.☆5,579Updated this week
- AIRA-dojo: a framework for developing and evaluating AI research agents☆154Apr 14, 2026Updated 3 months ago
- MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.☆592Jun 23, 2026Updated 3 weeks ago
- Search, understand, reproduce, and improve an idea with ease☆1,212Updated this week
- An agent benchmark with tasks in a simulated software company.☆748Nov 17, 2025Updated 8 months ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,472Jul 11, 2026Updated last week
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆708Jul 29, 2025Updated 11 months ago
- Post-training with Tinker☆3,869Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhi…☆815May 30, 2026Updated last month
- AllenAI's post-training codebase☆3,803Updated this week
- Ideas for projects related to Tinker☆191Nov 6, 2025Updated 8 months ago
- τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains☆1,631Updated this week
- 🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.☆667Jan 29, 2026Updated 5 months ago
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning☆199Mar 27, 2026Updated 3 months ago
- verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in…☆2,140Jun 9, 2026Updated last month