Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos (CVPR 2026)
☆26Dec 16, 2025Updated 8 months ago
Alternatives and similar repositories for VIPA-VLA
Users that are interested in VIPA-VLA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models (ECCV 2026)☆23Jul 2, 2026Updated last month
- Tackling Non-Stationarity in Reinforcement Learning via Causal-Origin Representation (ICML 2024)☆12Aug 9, 2024Updated 2 years ago
- AdaRefiner: Refining Decisions of Language Models with Adaptive Feedback (NAACL 2024)☆20Aug 9, 2024Updated 2 years ago
- Being-VL-0.5: Unified Multimodal Understanding via Byte-Pair Visual Encoding (ICCV 2025, Highlight)☆54Dec 22, 2025Updated 8 months ago
- A simple tool to help get information in NKU-EAMIS(NKU Education Affairs Management Information System).☆10Jul 27, 2020Updated 6 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- 一款基于DQN算法的牌类游戏AI框架 / An AI framework for card games based on DQN algorithm☆13Jul 25, 2024Updated 2 years ago
- Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models (ECCV 2026)☆38Jun 30, 2026Updated 2 months ago
- Kinetix simulation for Legato: Learning Native Continuation for Action Chunking Flow Policies (RSS 2026)☆28Jun 3, 2026Updated 2 months ago
- Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills☆64Jun 19, 2025Updated last year
- Being-H is BeingBeyond's family of human-centric embodied foundation models.☆1,131Updated this week
- ☆35Updated this week
- Awesome paper for multi-modal llm with grounding ability☆21Oct 11, 2025Updated 10 months ago
- My Blog (https://www.zhangwp.com).☆30Jan 11, 2024Updated 2 years ago
- [ICML 2026] This repo is the official implementation of "LangForce : Bayesian Decomposition of Vision Language Action Models via Latent …☆77Jul 29, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Pi0-VLA Repository of "MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies"☆28Mar 9, 2026Updated 5 months ago
- The repository provides code for EgoMAN model and dataset creation scripts.☆33Dec 31, 2025Updated 7 months ago
- [ICRA 2026] VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos☆483Jun 12, 2026Updated 2 months ago
- ☆39Mar 8, 2026Updated 5 months ago
- ChatEMG: Synthetic Data Generation to Control a Robotic Hand Orthosis for Stroke☆12Jul 2, 2024Updated 2 years ago
- [ECCV 2026] Official implementation of "RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics"☆83Jun 18, 2026Updated 2 months ago
- Official Repository of "Transcrib3D: 3D Referring Expression Resolution through Large Language Models" accepted at IROS 2024☆13Mar 30, 2026Updated 5 months ago
- Matlab source code of the paper "Y. Cui, D. Wu* and J. Huang*, "Optimize TSK Fuzzy Systems for Classification Problems: Mini-Batch Gradie…☆15Feb 7, 2021Updated 5 years ago
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos☆339Apr 16, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos (ICML 2026)☆59May 4, 2026Updated 3 months ago
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model (ICCV 2025)☆37Sep 4, 2025Updated 11 months ago
- ☆10Nov 30, 2022Updated 3 years ago
- Official implementation of Spatial-Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model [ICLR2026]☆281Jul 7, 2026Updated last month
- Implementation of "HumanReg: Self-supervised Non-rigid Registration of Sparse Human Point Cloud" (3DV 2024)☆15Oct 26, 2024Updated last year
- SGAP-Net: Semantic-Guided Attentive Prototypes Network for Few-Shot Human-Object Interaction Recognition, AAAI2020.☆14Dec 15, 2020Updated 5 years ago
- From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots (NeurIPS 2025, Spotlight)☆78Sep 28, 2025Updated 11 months ago
- [ICML 2026] Orienting Latent Actions for Video World Modeling☆121Apr 20, 2026Updated 4 months ago
- Human-centric environment representations from egocentric video☆15Feb 5, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Learning Precise Affordances from Egocentric Videos for Robotic Manipulation (ICCV 2025)☆26Jan 30, 2026Updated 7 months ago
- AAAI 2026 Oral☆19Dec 23, 2025Updated 8 months ago
- Implementation of the Mesh-VQVAE of "VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space" - ECCV 2024☆19Oct 30, 2024Updated last year
- Official implementation for AAAI-26 paper: "Force-Aware 3D Contact Modeling for Stable Grasp Generation"☆15Mar 13, 2026Updated 5 months ago
- DemoGrasp: Universal Dexterous Grasping from a Single Demonstration (ICLR 2026)☆87Feb 14, 2026Updated 6 months ago
- Human-in-the-loop Online Rejection Sampling for Robotic Manipulation☆27Nov 3, 2025Updated 9 months ago
- Reproduction of "How Does Batch Normalization Help Optimization?" paper☆21Mar 28, 2019Updated 7 years ago