Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.
β60Jul 14, 2026Updated 2 months ago
Alternatives and similar repositories for Terrarium
Users that are interested in Terrarium are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π¦ ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agentsβ124May 28, 2026Updated 4 months ago
- A unified framework for vision-language environments with Gymnasium-compatible interfaceβ38Mar 17, 2026Updated 6 months ago
- β59Apr 13, 2026Updated 5 months ago
- πͺ Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedbackβ26Jan 29, 2026Updated 8 months ago
- β44Jun 30, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- MM-Eureka V0 also called R1-Multimodal-Journey, Latest version is in MM-Eurekaβ325Jun 21, 2025Updated last year
- Revisiting Mid-training in the Era of Reinforcement Learning Scalingβ188Jul 23, 2025Updated last year
- β15Feb 21, 2024Updated 2 years ago
- Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?β20Mar 9, 2025Updated last year
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Spaceβ17Apr 16, 2026Updated 5 months ago
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervisionβ30May 26, 2025Updated last year
- β48May 12, 2026Updated 4 months ago
- β59Sep 2, 2024Updated 2 years ago
- Benchmarking Language Agents Under Controllable and Extreme Context Growthβ60Apr 29, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- The raw UserRL repo under constructionβ122Jun 2, 2026Updated 4 months ago
- [NeurIPS 2025 Spotlight] FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocitiesβ78Dec 21, 2025Updated 9 months ago
- εΌζΊζθͺε·± β A Claude Code skill trained on Flood Sung's entire Zhihu corpus (152 articles + 178 pins + 254 answers). Fork it to open-source yβ¦β25Apr 21, 2026Updated 5 months ago
- Various test models in WNNX format. It can view with `pip install wnetron && wnetron`β12Jun 22, 2022Updated 4 years ago
- reasoning-from-scratchηδΈζηΏ»θ―ηζ¬β66Sep 15, 2026Updated 2 weeks ago
- β22Nov 17, 2025Updated 10 months ago
- Official repo for "SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization"β29Mar 24, 2026Updated 6 months ago
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentationβ112Sep 18, 2025Updated last year
- Code for "Language Models Can Learn from Verbal Feedback Without Scalar Rewards"β65Jan 5, 2026Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.β169Aug 30, 2026Updated last month
- Code and Data for "Language Modeling with Editable External Knowledge"β39Jun 19, 2024Updated 2 years ago
- An Agent-native Obsidian wiki + CLAUDE.md skeleton that lets Agent know your project as well as you do.β29Apr 5, 2026Updated 5 months ago
- A full-stack online music app, developed using MERN stack (React, Express.js, MongoDB) and Electron. Libraries including Tailwind CSS, Reβ¦β10Jul 2, 2024Updated 2 years ago
- The official github repo for MixEval-X, the first any-to-any, real-world benchmark.β17Feb 15, 2025Updated last year
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelismβ18Nov 18, 2025Updated 10 months ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoningβ406Mar 30, 2026Updated 6 months ago
- β17Aug 2, 2023Updated 3 years ago
- Synchronizing Claude Code conversations across machinesβ17Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Templates and examples for ACL and EMNLP conference posters.β15Oct 5, 2024Updated last year
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention β¦β33Apr 2, 2026Updated 6 months ago
- The official github repo for "Training Optimal Large Diffusion Language Models", the first-ever large-scale diffusion language models scaβ¦β46Nov 6, 2025Updated 10 months ago
- Official Implementation of "Simulating Environments with Reasoning Models for Agent Training"β68Feb 18, 2026Updated 7 months ago
- Collections of RLxLM experiments using minimal codesβ14Feb 17, 2025Updated last year
- [EMNLP2026] Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forwardβ61Nov 27, 2025Updated 10 months ago
- An agent framework for building and evaluating general digital agents.β44Apr 21, 2026Updated 5 months ago