Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.
β59Jul 14, 2026Updated last month
Alternatives and similar repositories for Terrarium
Users that are interested in Terrarium are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π¦ ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agentsβ124May 28, 2026Updated 2 months ago
- A unified framework for vision-language environments with Gymnasium-compatible interfaceβ38Mar 17, 2026Updated 5 months ago
- β59Apr 13, 2026Updated 4 months ago
- πͺ Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedbackβ23Jan 29, 2026Updated 6 months ago
- [NeurIPS'2025] "OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis"β31Dec 4, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- β41Jun 30, 2026Updated last month
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmarkβ27Apr 13, 2026Updated 4 months ago
- MM-Eureka V0 also called R1-Multimodal-Journey, Latest version is in MM-Eurekaβ325Jun 21, 2025Updated last year
- Revisiting Mid-training in the Era of Reinforcement Learning Scalingβ188Jul 23, 2025Updated last year
- β15Feb 21, 2024Updated 2 years ago
- Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?β20Mar 9, 2025Updated last year
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Spaceβ17Apr 16, 2026Updated 4 months ago
- This is an official code for the paper: TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generationβ29Mar 26, 2026Updated 4 months ago
- β41May 12, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervisionβ30May 26, 2025Updated last year
- β59Sep 2, 2024Updated last year
- β19Nov 17, 2025Updated 9 months ago
- β50Updated this week
- The raw UserRL repo under constructionβ117Jun 2, 2026Updated 2 months ago
- [NeurIPS 2025 Spotlight] FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocitiesβ78Dec 21, 2025Updated 8 months ago
- εΌζΊζθͺε·± β A Claude Code skill trained on Flood Sung's entire Zhihu corpus (152 articles + 178 pins + 254 answers). Fork it to open-source yβ¦β25Apr 21, 2026Updated 4 months ago
- Various test models in WNNX format. It can view with `pip install wnetron && wnetron`β12Jun 22, 2022Updated 4 years ago
- [ACL 2025 Findings] Text2World: Benchmarking Large Language Models for Symbolic World Model Generationβ29Feb 25, 2025Updated last year
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official repo for "SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization"β29Mar 24, 2026Updated 5 months ago
- Notion LifeOS PARA system β agent skill for Claude Code, OpenClaw, Codex and moreβ23Mar 24, 2026Updated 5 months ago
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentationβ112Sep 18, 2025Updated 11 months ago
- Code for "Language Models Can Learn from Verbal Feedback Without Scalar Rewards"β65Jan 5, 2026Updated 7 months ago
- A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.β161Jun 17, 2026Updated 2 months ago
- Code and Data for "Language Modeling with Editable External Knowledge"β38Jun 19, 2024Updated 2 years ago
- β15Oct 13, 2025Updated 10 months ago
- A full-stack online music app, developed using MERN stack (React, Express.js, MongoDB) and Electron. Libraries including Tailwind CSS, Reβ¦β10Jul 2, 2024Updated 2 years ago
- An Agent-native Obsidian wiki + CLAUDE.md skeleton that lets Agent know your project as well as you do.β29Apr 5, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Benchmarking Attention Mechanism in Vision Transformers.β20Oct 10, 2022Updated 3 years ago
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelismβ18Nov 18, 2025Updated 9 months ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoningβ404Mar 30, 2026Updated 4 months ago
- Synchronizing Claude Code conversations across machinesβ17Aug 14, 2026Updated last week
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention β¦β31Apr 2, 2026Updated 4 months ago
- The official github repo for "Training Optimal Large Diffusion Language Models", the first-ever large-scale diffusion language models scaβ¦β46Nov 6, 2025Updated 9 months ago
- Official Implementation of "Simulating Environments with Reasoning Models for Agent Training"β67Feb 18, 2026Updated 6 months ago