Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
☆539Aug 10, 2026Updated this week
Alternatives and similar repositories for meta-agents-research-environments
Users that are interested in meta-agents-research-environments are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Gym for Agentic LLMs☆504Jan 21, 2026Updated 6 months ago
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,138Updated this week
- The raw UserRL repo under construction☆115Jun 2, 2026Updated 2 months ago
- ☆311Jul 1, 2026Updated last month
- An interface library for RL post training with environments.☆2,490Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Our library for RL environments + evals☆4,485Updated this week
- Research code artifacts for Code World Model (CWM) including inference tools, reproducibility, and documentation.☆887Jul 17, 2026Updated 3 weeks ago
- MCP Atlas☆141Updated this week
- Framework for evaluating and improving agents☆4,076Updated this week
- Democratizing Reinforcement Learning for LLMs☆5,774Updated this week
- A version of verl to support diverse tool use [TMLR 2026]☆1,030Jul 15, 2026Updated 3 weeks ago
- slime is an LLM post-training framework for RL Scaling.☆7,832Updated this week
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,477Apr 17, 2026Updated 3 months ago
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆72Jan 28, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆494Updated this week
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆449Updated this week
- ☆1,175Jan 10, 2026Updated 7 months ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,681Apr 24, 2026Updated 3 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆22,900Updated this week
- RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.☆2,765Jul 24, 2026Updated 2 weeks ago
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆1,945Updated this week
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆145May 6, 2026Updated 3 months ago
- Harness for running and evaluating AI agents against RL environments☆232Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Agentic RL Training at Scale☆1,878Updated this week
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.☆5,653Updated this week
- AIRA-dojo: a framework for developing and evaluating AI research agents☆157Apr 14, 2026Updated 3 months ago
- MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.☆592Jun 23, 2026Updated last month
- Search, understand, reproduce, and improve an idea with ease☆1,218Updated this week
- An agent benchmark with tasks in a simulated software company.☆758Nov 17, 2025Updated 8 months ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,534Jul 11, 2026Updated last month
- Post-training with Tinker☆4,005Updated this week
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆720Jul 29, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhi…☆828May 30, 2026Updated 2 months ago
- AllenAI's post-training codebase☆3,826Updated this week
- Ideas for projects related to Tinker☆192Nov 6, 2025Updated 9 months ago
- τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains☆1,770Updated this week
- 🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.☆667Jan 29, 2026Updated 6 months ago
- MLGym A New Framework and Benchmark for Advancing AI Research Agents☆616Aug 10, 2025Updated last year
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning☆201Mar 27, 2026Updated 4 months ago