☆49Aug 26, 2025Updated 10 months ago
Alternatives and similar repositories for embodied-videoagent
Users that are interested in embodied-videoagent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆55Oct 3, 2024Updated last year
- ☆13Nov 5, 2024Updated last year
- ☆266Aug 6, 2025Updated 11 months ago
- ☆15Mar 16, 2026Updated 4 months ago
- [ICLR'26] This repository is the implementation of "3D Aware Region Prompted Vision Language Model"☆28Feb 19, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆18Dec 1, 2025Updated 7 months ago
- VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation☆31Nov 3, 2025Updated 8 months ago
- ☆30Jun 19, 2024Updated 2 years ago
- [CVPR25] SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse Inputs☆20Aug 27, 2025Updated 10 months ago
- [CVPR 2025] Source codes for the paper "3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning"☆266Oct 2, 2025Updated 9 months ago
- ☆19Aug 7, 2025Updated 11 months ago
- MAT: Multi-modal Agent Tuning 🔥 ICLR 2025 (Spotlight)☆97Dec 18, 2025Updated 7 months ago
- FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation Models, ICCV 2023☆13Jul 13, 2024Updated 2 years ago
- DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric Voxelization (ICLR2024) & DynaVol-S: Dynamic Scene Understanding…☆21Apr 10, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICLR 2026] SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes☆27Mar 22, 2026Updated 3 months ago
- We present SceneFactor, a diffusion-based approach for large-scale 3D scene generation that enables controllable generation and effortles…☆106Apr 29, 2026Updated 2 months ago
- [CVPR 2025] Beacon3D: Object-centric Evaluation for 3D Grounding-QA☆28Nov 25, 2025Updated 7 months ago
- PyTorch implementation of the paper: CASAGPT: Cuboid Arrangement and Scene Assembly for Interior Design [CVPR 2025]☆15Apr 5, 2025Updated last year
- [CVPR 2026] Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation☆21Jun 11, 2026Updated last month
- ☆14Nov 14, 2023Updated 2 years ago
- ☆19Jul 20, 2024Updated 2 years ago
- Learning Precise Affordances from Egocentric Videos for Robotic Manipulation (ICCV 2025)☆26Jan 30, 2026Updated 5 months ago
- Code for the paper "Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration"☆56May 22, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code release for 'Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs' (NeurIPS 2025)☆31Oct 28, 2025Updated 8 months ago
- A video question answering dataset that focuses on the dynamics properties of objects (velocity, acceleration) and their collisions withi…☆20Apr 23, 2025Updated last year
- [IROS'25] Official implementation of the paper FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction☆17Oct 11, 2025Updated 9 months ago
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs☆69Mar 22, 2026Updated 3 months ago
- [RA-L 2026] Official implemetation of the paper "FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipu…☆19Jan 19, 2026Updated 6 months ago
- [ICLR 2026 Oral] MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning☆56Mar 4, 2026Updated 4 months ago
- ☆68Dec 3, 2025Updated 7 months ago
- [ICCV'25] T2 -VLM: Training-Free Generation of Temporally Consistent Rewards from VLMs☆16Jul 8, 2025Updated last year
- [WACV2023] Intention-Conditioned Long-Term Human Egocentric Action Forecasting @ EGO4D Challenge 2022☆14Sep 3, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Meta-Memory: Retrieving and Integrating Semantic-Spatial Memories for Robot Spatial Reasoning☆16Nov 26, 2025Updated 7 months ago
- CVPR 2026 - MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation☆64Mar 23, 2026Updated 3 months ago
- This is the official code of VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding (ECCV 2024)☆320Dec 5, 2024Updated last year
- Nav-R1: Reasoning and Navigation in Embodied Scenes☆128Oct 31, 2025Updated 8 months ago
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics☆23Nov 18, 2025Updated 8 months ago
- [CVPR 2026 (Oral)] MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping☆156May 27, 2026Updated last month
- ☆16Jan 27, 2026Updated 5 months ago