[NeurIPS 2024] Official Implementation for Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
☆104Jun 17, 2025Updated last year
Alternatives and similar repositories for NeurIPS24-Optimus-1
Users that are interested in NeurIPS24-Optimus-1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Paper List of Minecraft Agents☆71May 24, 2026Updated 4 months ago
- [CVPR 2025] Official Implementation for Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy☆28Jun 17, 2025Updated last year
- [ICML 2024] Official repository of ICML 2024 - RoboMP2: A Robotic Multimodal Perception-Planning Framework with Multimodal Large Language…☆12Apr 4, 2026Updated 6 months ago
- Official repository of the "Fine-grained Key-Value Memory Enhanced Predictor for Video Representation Learning" (ACM MM 2023)☆23Jul 11, 2024Updated 2 years ago
- Official Implementation for Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimiza…☆74Apr 14, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2026] HiconAgent: History Context-aware Policy Optimization for GUI Agents☆33Mar 9, 2026Updated 7 months ago
- [ECCV 2024] STEVE in Minecraft is for See and Think: Embodied Agent in Virtual Environment☆42Dec 27, 2023Updated 2 years ago
- Official implementation of paper "ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting" (CVPR'25)☆48Apr 13, 2025Updated last year
- Official repository of the “Mask Again: Masked Knowledge Distillation for Masked Video Modeling” (ACM MM 2023)☆27Jul 11, 2024Updated 2 years ago
- STEVE-1: A Generative Model for Text-to-Behavior in Minecraft☆228Jun 4, 2024Updated 2 years ago
- This repo is a live list of papers on game playing and large multimodality model - "A Survey on Game Playing Agents and Large Models: Met…☆162Sep 3, 2024Updated 2 years ago
- [ACL 2026 main] PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records☆30Apr 11, 2026Updated 5 months ago
- Official code for the paper: WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents☆70Dec 3, 2025Updated 10 months ago
- [IROS'25 Oral & NeurIPSw'24] Official implementation of "MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simula…☆104Jun 16, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆51Dec 11, 2023Updated 2 years ago
- HEtero-Assists Distillation for Heterogeneous Object Detectors☆10Jul 3, 2023Updated 3 years ago
- [CVPR2024] This is the official implement of MP5☆108Jun 30, 2024Updated 2 years ago
- [ACM MM 2025] PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning☆19Jun 6, 2026Updated 4 months ago
- [NeurIPS 2024] MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models☆87Dec 27, 2025Updated 9 months ago
- JARVIS-1: Open-world Multi-task Agents with Memory-Augmented Multimodal Language Models☆418Apr 8, 2024Updated 2 years ago
- ☆60Oct 21, 2025Updated 11 months ago
- GROOT: Learning to Follow Instructions by Watching Gameplay Videos (ICLR'24, Spotlight)☆71Dec 18, 2023Updated 2 years ago
- Support finetuning GLM4v with zero2☆16Jun 29, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆47Jun 11, 2025Updated last year
- [CVPR 2025] LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant☆29Dec 2, 2025Updated 10 months ago
- [ICLR 2025] Official Implementation for 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3…☆19Apr 7, 2026Updated 6 months ago
- ☆31Jun 25, 2024Updated 2 years ago
- Odyssey: Empowering Minecraft Agents with Open-World Skills☆409Oct 22, 2025Updated 11 months ago
- [NeurIPS 2024] The official implementation of "Image Copy Detection for Diffusion Models"☆18Oct 1, 2024Updated 2 years ago
- ☆31Apr 30, 2026Updated 5 months ago
- MAKGED is the first multi-agent framework for collaborative error detection in knowledge graphs.☆31Jul 20, 2025Updated last year
- [ACM MM 2025] Official repository of "EmoSym: A Symbiotic Framework for Unified Emotional Understanding and Generation via Latent Reasoni…☆32May 6, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation of the DECKARD Agent from the paper "Do Embodied Agents Dream of Pixelated Sheep?"☆95May 23, 2023Updated 3 years ago
- ☆46Dec 30, 2024Updated last year
- [ICCV 2025] GUIOdyssey is a comprehensive dataset for training and evaluating cross-app navigation agents. GUIOdyssey consists of 8,834 e…☆162Jan 3, 2026Updated 9 months ago
- ☆101Jun 12, 2024Updated 2 years ago
- [NeurIPS 2025] CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification☆192Jun 17, 2026Updated 3 months ago
- Simulating Large-Scale Multi-Agent Interactions with Limited Multimodal Senses and Physical Needs☆113Sep 30, 2025Updated last year
- F1: A Vision Language Action Model Bridging Understanding and Generation to Actions☆199Jan 2, 2026Updated 9 months ago