[ACM MM'26 Oral] MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
☆68May 14, 2026Updated 4 months ago
Alternatives and similar repositories for MMaDA-VLA
Users that are interested in MMaDA-VLA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'25] SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning☆42Oct 14, 2025Updated 11 months ago
- ActionCodec: What Makes for Good Action Tokenizers☆74Mar 1, 2026Updated 7 months ago
- [ICLR 2026] Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining☆34Apr 26, 2026Updated 5 months ago
- Official implementation of FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment☆56Mar 24, 2026Updated 6 months ago
- [NeurIPS 2026] FASTER: Rethinking Real-Time Flow VLAs☆157May 14, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2026] ResVLA: From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges☆29Jun 1, 2026Updated 4 months ago
- LAP: Language-Action Pre-Training Enables Zero-Shot Cross Embodiment Transfer☆173May 20, 2026Updated 4 months ago
- Dream-VL and Dream-VLA, a diffusion VLM and a diffusion VLA.☆116Jan 14, 2026Updated 8 months ago
- [RSS 2026] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion☆321May 26, 2026Updated 4 months ago
- 🔥 The first open-sourced diffusion vision-langauge-action model. [ICLR 2026]☆186Mar 12, 2026Updated 6 months ago
- Official implementation of Spatial-Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model [ICLR2026]☆290Jul 7, 2026Updated 2 months ago
- Reshaping Action Error Distributions for Reliable Vision-Language-Action Models☆17Feb 5, 2026Updated 7 months ago
- ☆50May 12, 2026Updated 4 months ago
- Code for Xiaomi-Robotics-1☆714Sep 9, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆34Jun 7, 2026Updated 3 months ago
- [CVPR 2026] HiF-VLA: An efficient, bidirectional spatiotemporal expansion Vision-Language-Action Model☆78Mar 11, 2026Updated 6 months ago
- [ICLR 2026] Unified Vision-Language-Action Model☆322Oct 15, 2025Updated 11 months ago
- RynnVLA-002: A Unified Vision-Language-Action and World Model☆1,135Dec 2, 2025Updated 10 months ago
- Cosmos Policy☆879Jan 23, 2026Updated 8 months ago
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation☆60Aug 10, 2026Updated last month
- [RSS 2026] Causal video-action world model for generalist robot control☆1,922Jul 9, 2026Updated 2 months ago
- Self-CorrectingVLA:OnlineActionRefinementviaSparseWorldImagination☆27Apr 14, 2026Updated 5 months ago
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing☆3,753Sep 23, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆50Jun 30, 2026Updated 3 months ago
- Official code of Motus: A Unified Latent Action World Model☆1,299Jan 5, 2026Updated 8 months ago
- The official implementation of MaskGRPO: Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models. (ICLR 2026, arxiv…☆19Jan 27, 2026Updated 8 months ago
- Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals☆2,680Apr 19, 2026Updated 5 months ago
- An all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.☆720Updated this week
- Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?☆1,539Aug 20, 2026Updated last month
- This Repository will be used by the Concordia Shanghai Swarm Team Club to share & work on different programs for the DJI TT Swarm Kit☆10Dec 19, 2021Updated 4 years ago
- [ICRA'25] Official code repository of "QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning"☆23Jun 25, 2026Updated 3 months ago
- LIBERO-PRO is the official repository of the LIBERO-PRO — an evaluation extension of the original LIBERO benchmark☆331Jul 17, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [CVPR'26] QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models☆47Mar 25, 2026Updated 6 months ago
- InternVLA-A1: Unifying Understanding, Generation, and Action for Robotic Manipulation☆563Sep 14, 2026Updated 2 weeks ago
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success☆1,407Sep 9, 2025Updated last year
- Implementation of "SimVLA: A Simple VLA Baseline for Robotic Manipulation"☆147Feb 26, 2026Updated 7 months ago
- Official implementation of ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.☆279Apr 1, 2026Updated 6 months ago
- GeRM: A Generalist Robotic Model with Mixture-of-Experts for Quadruped Robot https://songwxuan.github.io/GeRM/☆38Apr 29, 2025Updated last year
- RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation☆90Jul 16, 2026Updated 2 months ago