Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos (CVPR 2026)
☆26Dec 16, 2025Updated 7 months ago
Alternatives and similar repositories for VIPA-VLA
Users that are interested in VIPA-VLA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models (ECCV 2026)☆22Jul 2, 2026Updated last month
- AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback (NAACL 2024)☆19Aug 9, 2024Updated 2 years ago
- Being-VL-0.5: Unified Multimodal Understanding via Byte-Pair Visual Encoding (ICCV 2025, Highlight)☆54Dec 22, 2025Updated 7 months ago
- Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models (ECCV 2026)☆38Jun 30, 2026Updated last month
- UniTacHand: Unified Spatio-Tactile Representation for Human-to-Dexterous-Hand Skill Transfer☆26Dec 25, 2025Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Kinetix simulation for Legato: Learning Native Continuation for Action Chunking Flow Policies (RSS 2026)☆24Jun 3, 2026Updated 2 months ago
- Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills☆63Jun 19, 2025Updated last year
- Being-H is BeingBeyond's family of human-centric embodied foundation models.☆1,117Updated this week
- Awesome paper for multi-modal llm with grounding ability☆21Oct 11, 2025Updated 9 months ago
- [ICML 2026] This repo is the official implementation of "LangForce : Bayesian Decomposition of Vision Language Action Models via Latent …☆75Jul 29, 2026Updated last week
- Pi0-VLA Repository of "MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies"☆28Mar 9, 2026Updated 5 months ago
- The repository provides code for EgoMAN model and dataset creation scripts.☆32Dec 31, 2025Updated 7 months ago
- [ICRA 2026] VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos☆469Jun 12, 2026Updated last month
- ☆39Mar 8, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ECCV 2026] Official implementation of "RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics"☆82Jun 18, 2026Updated last month
- ChatEMG: Synthetic Data Generation to Control a Robotic Hand Orthosis for Stroke☆12Jul 2, 2024Updated 2 years ago
- Official Repository of "Transcrib3D: 3D Referring Expression Resolution through Large Language Models" accepted at IROS 2024☆13Mar 30, 2026Updated 4 months ago
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos☆317Apr 16, 2026Updated 3 months ago
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos (ICML 2026)☆56May 4, 2026Updated 3 months ago
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model (ICCV 2025)☆37Sep 4, 2025Updated 11 months ago
- Official implementation of Spatial-Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model [ICLR2026]☆279Jul 7, 2026Updated last month
- Implementation of "HumanReg: Self-supervised Non-rigid Registration of Sparse Human Point Cloud" (3DV 2024)☆15Oct 26, 2024Updated last year
- SGAP-Net: Semantic-Guided Attentive Prototypes Network for Few-Shot Human-Object Interaction Recognition, AAAI2020.☆14Dec 15, 2020Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots (NeurIPS 2025, Spotlight)☆79Sep 28, 2025Updated 10 months ago
- [ICML 2026] Orienting Latent Actions for Video World Modeling☆120Apr 20, 2026Updated 3 months ago
- Human-centric environment representations from egocentric video☆15Feb 5, 2026Updated 6 months ago
- Learning Precise Affordances from Egocentric Videos for Robotic Manipulation (ICCV 2025)☆26Jan 30, 2026Updated 6 months ago
- AAAI 2026 Oral☆19Dec 23, 2025Updated 7 months ago
- Implementation of the Mesh-VQVAE of "VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space" - ECCV 2024☆18Oct 30, 2024Updated last year
- Official implementation for AAAI-26 paper: "Force-Aware 3D Contact Modeling for Stable Grasp Generation"☆15Mar 13, 2026Updated 4 months ago
- DemoGrasp: Universal Dexterous Grasping from a Single Demonstration (ICLR 2026)☆84Feb 14, 2026Updated 5 months ago
- Human-in-the-loop Online Rejection Sampling for Robotic Manipulation☆27Nov 3, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Cascade xDAWN EEGNet for ERP detection☆19Apr 27, 2024Updated 2 years ago
- Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces☆86Jun 6, 2025Updated last year
- EO: Open-source Unified Embodied Foundation Model Series☆59Jan 15, 2026Updated 6 months ago
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models☆88Dec 17, 2025Updated 7 months ago
- CoRL25-"AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies"☆50Aug 15, 2025Updated 11 months ago
- Getting Started in Imitation Learning☆13Mar 3, 2025Updated last year
- EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning (ICLR 2026)☆33Jan 29, 2026Updated 6 months ago