Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.
β54Jul 14, 2026Updated 3 weeks ago
Alternatives and similar repositories for Terrarium
Users that are interested in Terrarium are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π¦ ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agentsβ120May 28, 2026Updated 2 months ago
- A unified framework for vision-language environments with Gymnasium-compatible interfaceβ35Mar 17, 2026Updated 4 months ago
- β59Apr 13, 2026Updated 3 months ago
- πͺ Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedbackβ23Jan 29, 2026Updated 6 months ago
- [NeurIPS'2025] "OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis"β30Dec 4, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β33Jun 30, 2026Updated last month
- Searching a High Performance Feature Extractor for Text Recognition Network. TPAMI 2022β13Nov 25, 2022Updated 3 years ago
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmarkβ26Apr 13, 2026Updated 3 months ago
- MM-Eureka V0 also called R1-Multimodal-Journey, Latest version is in MM-Eurekaβ325Jun 21, 2025Updated last year
- β15Feb 21, 2024Updated 2 years ago
- Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?β20Mar 9, 2025Updated last year
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Spaceβ17Apr 16, 2026Updated 3 months ago
- This is an official code for the paper: TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generationβ28Mar 26, 2026Updated 4 months ago
- β41May 12, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervisionβ30May 26, 2025Updated last year
- β59Sep 2, 2024Updated last year
- Benchmarking Language Agents Under Controllable and Extreme Context Growthβ51Apr 29, 2026Updated 3 months ago
- β19Nov 17, 2025Updated 8 months ago
- β50Updated this week
- [NeurIPS 2025 Spotlight] FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocitiesβ77Dec 21, 2025Updated 7 months ago
- εΌζΊζθͺε·± β A Claude Code skill trained on Flood Sung's entire Zhihu corpus (152 articles + 178 pins + 254 answers). Fork it to open-source yβ¦β25Apr 21, 2026Updated 3 months ago
- reasoning-from-scratchηδΈζηΏ»θ―ηζ¬β52Dec 5, 2025Updated 7 months ago
- [ACL 2025 Findings] Text2World: Benchmarking Large Language Models for Symbolic World Model Generationβ29Feb 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official repo for "SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization"β28Mar 24, 2026Updated 4 months ago
- Notion LifeOS PARA system β agent skill for Claude Code, OpenClaw, Codex and moreβ23Mar 24, 2026Updated 4 months ago
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentationβ112Sep 18, 2025Updated 10 months ago
- Code for "Language Models Can Learn from Verbal Feedback Without Scalar Rewards"β65Jan 5, 2026Updated 6 months ago
- A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.β158Jun 17, 2026Updated last month
- Code and Data for "Language Modeling with Editable External Knowledge"β37Jun 19, 2024Updated 2 years ago
- β15Oct 13, 2025Updated 9 months ago
- A full-stack online music app, developed using MERN stack (React, Express.js, MongoDB) and Electron. Libraries including Tailwind CSS, Reβ¦β10Jul 2, 2024Updated 2 years ago
- An Agent-native Obsidian wiki + CLAUDE.md skeleton that lets Agent know your project as well as you do.β29Apr 5, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The official github repo for MixEval-X, the first any-to-any, real-world benchmark.β17Feb 15, 2025Updated last year
- Benchmarking Attention Mechanism in Vision Transformers.β20Oct 10, 2022Updated 3 years ago
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelismβ18Nov 18, 2025Updated 8 months ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoningβ403Mar 30, 2026Updated 4 months ago
- β17Aug 2, 2023Updated 3 years ago
- Synchronizing Claude Code conversations across machinesβ16Jul 23, 2026Updated last week
- Templates and examples for ACL and EMNLP conference posters.β15Oct 5, 2024Updated last year