π» SETA: Scaling Environments for Terminal Agents
β124Jul 17, 2026Updated this week
Alternatives and similar repositories for seta
Users that are interested in seta are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π» SETA: Scaling Environments for Terminal Agents - Environmentsβ142Feb 16, 2026Updated 5 months ago
- β134Mar 31, 2026Updated 3 months ago
- GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's Tβ¦β395Aug 24, 2025Updated 10 months ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL trainingβ70Jul 28, 2025Updated 11 months ago
- SWE-Bench-plus-plusβ25Feb 5, 2026Updated 5 months ago
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Framework for evaluating and improving agentsβ3,320Updated this week
- Harness for running and evaluating AI agents against RL environmentsβ217Updated this week
- A Python SDK for Open Reward Standard servers and clientsβ17Mar 24, 2026Updated 3 months ago
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Executionβ430Updated this week
- Official Implementation of "Simulating Environments with Reasoning Models for Agent Training"β65Feb 18, 2026Updated 5 months ago
- Toolathlon-Gym for testing AI agents real-world tool-use capabilities across diverse MCP servers.β138Apr 2, 2026Updated 3 months ago
- β119Apr 1, 2026Updated 3 months ago
- β52May 26, 2026Updated last month
- SkyRL: A Modular Full-stack RL Library for LLMsβ2,081Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Data recipes and robust infrastructure for training AI agentsβ260Updated this week
- π Loong: Synthesize Long CoTs at Scale through Verifiers.β506Jul 10, 2026Updated last week
- [EMNLP 2024 Main] Official implementation of the paper "The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Languaβ¦β13Nov 11, 2024Updated last year
- [COLM 2025] Official repository for R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agentsβ307Jul 13, 2025Updated last year
- [NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agentsβ709Jul 13, 2026Updated last week
- Nemotron-CORTEXA is an open-source software engineering agent that fixes GitHub issues.β25Aug 7, 2025Updated 11 months ago
- COLM2026β36Jul 9, 2026Updated last week
- A compact high-signal benchmark for evaluating frontier agentsβ18Updated this week
- β16Jul 11, 2026Updated last week
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenariosβ586Jun 12, 2026Updated last month
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.β5,575Updated this week
- Agentic RL on Any Harness at Scaleβ688Updated this week
- open source SWE-Atlasβ55Updated this week
- Official implementation of Selective Entropy Regularization (SIREN), proposed by paper 'Rethinking Entropy Regularization in Large Reasonβ¦β32Dec 10, 2025Updated 7 months ago
- [COLM 2025] Official repo for Self-Steering Language Modelsβ26Aug 8, 2025Updated 11 months ago
- [FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarksβ183May 12, 2026Updated 2 months ago
- β36May 16, 2026Updated 2 months ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scaleβ47May 30, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Democratizing Reinforcement Learning for LLMsβ5,708Updated this week
- slime is an LLM post-training framework for RL Scaling.β7,551Updated this week
- A visual representation of Dijkstra's Algorithm using Libgdx.β16Dec 24, 2021Updated 4 years ago
- β13Jan 14, 2025Updated last year
- [NSDI'26] PolyRL is a reinforcement learning framework for LLM that harvest spot instances on the cloud to reduce cost.β19Mar 30, 2026Updated 3 months ago
- Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Frameworkβ321Jan 17, 2026Updated 6 months ago
- Concise Reasoning via Reinforcement Learningβ13Apr 16, 2025Updated last year