Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.
β60Jul 14, 2026Updated last month
Alternatives and similar repositories for Terrarium
Users that are interested in Terrarium are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π¦ ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agentsβ124May 28, 2026Updated 3 months ago
- β59Apr 13, 2026Updated 5 months ago
- πͺ Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedbackβ25Jan 29, 2026Updated 7 months ago
- [NeurIPS'2025] "OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis"β32Dec 4, 2025Updated 9 months ago
- β42Jun 30, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmarkβ27Apr 13, 2026Updated 5 months ago
- MM-Eureka V0 also called R1-Multimodal-Journey, Latest version is in MM-Eurekaβ325Jun 21, 2025Updated last year
- Revisiting Mid-training in the Era of Reinforcement Learning Scalingβ188Jul 23, 2025Updated last year
- β15Feb 21, 2024Updated 2 years ago
- Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?β20Mar 9, 2025Updated last year
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Spaceβ17Apr 16, 2026Updated 4 months ago
- This is an official code for the paper: TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generationβ28Mar 26, 2026Updated 5 months ago
- β43May 12, 2026Updated 4 months ago
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervisionβ30May 26, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Benchmarking Language Agents Under Controllable and Extreme Context Growthβ60Apr 29, 2026Updated 4 months ago
- UEval: A Benchmark for Unified Multimodal Generationβ26Apr 20, 2026Updated 4 months ago
- The raw UserRL repo under constructionβ120Jun 2, 2026Updated 3 months ago
- [NeurIPS 2025 Spotlight] FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocitiesβ78Dec 21, 2025Updated 8 months ago
- Various test models in WNNX format. It can view with `pip install wnetron && wnetron`β12Jun 22, 2022Updated 4 years ago
- [ACL 2025 Findings] Text2World: Benchmarking Large Language Models for Symbolic World Model Generationβ31Feb 25, 2025Updated last year
- β22Nov 17, 2025Updated 9 months ago
- Official repo for "SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization"β29Mar 24, 2026Updated 5 months ago
- Notion LifeOS PARA system β agent skill for Claude Code, OpenClaw, Codex and moreβ23Mar 24, 2026Updated 5 months ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentationβ112Sep 18, 2025Updated 11 months ago
- Code for "Language Models Can Learn from Verbal Feedback Without Scalar Rewards"β65Jan 5, 2026Updated 8 months ago
- A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.β164Aug 30, 2026Updated 2 weeks ago
- Code and Data for "Language Modeling with Editable External Knowledge"β39Jun 19, 2024Updated 2 years ago
- β15Oct 13, 2025Updated 11 months ago
- A full-stack online music app, developed using MERN stack (React, Express.js, MongoDB) and Electron. Libraries including Tailwind CSS, Reβ¦β10Jul 2, 2024Updated 2 years ago
- An Agent-native Obsidian wiki + CLAUDE.md skeleton that lets Agent know your project as well as you do.β29Apr 5, 2026Updated 5 months ago
- The official github repo for MixEval-X, the first any-to-any, real-world benchmark.β17Feb 15, 2025Updated last year
- Benchmarking Attention Mechanism in Vision Transformers.β20Oct 10, 2022Updated 3 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelismβ18Nov 18, 2025Updated 9 months ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoningβ405Mar 30, 2026Updated 5 months ago
- β17Aug 2, 2023Updated 3 years ago
- Templates and examples for ACL and EMNLP conference posters.β15Oct 5, 2024Updated last year
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention β¦β32Apr 2, 2026Updated 5 months ago
- The official github repo for "Training Optimal Large Diffusion Language Models", the first-ever large-scale diffusion language models scaβ¦β46Nov 6, 2025Updated 10 months ago
- Adversarial Training with Fast Gradient Projection Method against Synonym Substitution based Text Attacksβ24Dec 11, 2020Updated 5 years ago