Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
☆561Sep 30, 2026Updated last week
Alternatives and similar repositories for meta-agents-research-environments
Users that are interested in meta-agents-research-environments are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Gym for Agentic LLMs☆510Jan 21, 2026Updated 8 months ago
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,397Updated this week
- The raw UserRL repo under construction☆122Jun 2, 2026Updated 4 months ago
- ☆310Jul 1, 2026Updated 3 months ago
- An interface library for RL post training with environments.☆2,684Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Our library for RL environments + evals☆4,686Updated this week
- MCP Atlas☆158Oct 3, 2026Updated last week
- Research code artifacts for Code World Model (CWM) including inference tools, reproducibility, and documentation.☆910Jul 17, 2026Updated 2 months ago
- Framework for evaluating and improving agents☆5,951Updated this week
- Democratizing Reinforcement Learning for LLMs☆5,860Updated this week
- A version of verl to support diverse tool use [TMLR 2026]☆1,046Jul 15, 2026Updated 2 months ago
- slime is an LLM post-training framework for RL Scaling.☆8,615Updated this week
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,524Apr 17, 2026Updated 5 months ago
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆75Jan 28, 2026Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆588Oct 2, 2026Updated last week
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆491Aug 18, 2026Updated last month
- ☆1,193Jan 10, 2026Updated 9 months ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,769Apr 24, 2026Updated 5 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,805Updated this week
- Agent RL framework for LLM agents: multi-turn reinforcement learning with StarPO and reasoning-collapse diagnostics☆2,818Aug 23, 2026Updated last month
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆146May 6, 2026Updated 5 months ago
- Harness for running and evaluating AI agents against RL environments☆283Updated this week
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.☆5,823Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Agentic RL Training at Scale☆2,137Updated this week
- AIRA-dojo: a framework for developing and evaluating AI research agents☆174Apr 14, 2026Updated 5 months ago
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆3,075Updated this week
- MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.☆602Jun 23, 2026Updated 3 months ago
- Search, understand, reproduce, and improve an idea with ease☆1,241Updated this week
- An agent benchmark with tasks in a simulated software company.☆792Nov 17, 2025Updated 10 months ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,600Jul 11, 2026Updated 2 months ago
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆748Jul 29, 2025Updated last year
- Post-training with Tinker☆4,181Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhi…☆850May 30, 2026Updated 4 months ago
- AllenAI's post-training codebase☆3,882Updated this week
- Ideas for projects related to Tinker☆199Nov 6, 2025Updated 11 months ago
- 🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.☆671Jan 29, 2026Updated 8 months ago
- τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains☆2,198Updated this week
- MLGym A New Framework and Benchmark for Advancing AI Research Agents☆626Aug 10, 2025Updated last year
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning☆208Mar 27, 2026Updated 6 months ago