[ACM MM'26] MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
☆63May 14, 2026Updated 2 months ago
Alternatives and similar repositories for MMaDA-VLA
Users that are interested in MMaDA-VLA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'25] SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning☆40Oct 14, 2025Updated 9 months ago
- ActionCodec: What Makes for Good Action Tokenizers☆58Mar 1, 2026Updated 5 months ago
- [ICLR 2026] Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining☆31Apr 26, 2026Updated 3 months ago
- Official implementation of FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment☆55Mar 24, 2026Updated 4 months ago
- FASTER: Rethinking Real-Time Flow VLAs☆140May 14, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICML 2026] ResVLA: From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges☆28Jun 1, 2026Updated 2 months ago
- LAP: Language-Action Pre-Training Enables Zero-Shot Cross Embodiment Transfer☆160May 20, 2026Updated 2 months ago
- Dream-VL and Dream-VLA, a diffusion VLM and a diffusion VLA.☆114Jan 14, 2026Updated 6 months ago
- [RSS 2026] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion☆292May 26, 2026Updated 2 months ago
- 🔥 The first open-sourced diffusion vision-langauge-action model. [ICLR 2026]☆185Mar 12, 2026Updated 4 months ago
- Official implementation of Spatial-Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model [ICLR2026]☆272Jul 7, 2026Updated 3 weeks ago
- Reshaping Action Error Distributions for Reliable Vision-Language-Action Models☆17Feb 5, 2026Updated 5 months ago
- Code for Xiaomi-Robotics-1☆292Updated this week
- ☆49May 12, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆32Jun 7, 2026Updated last month
- [CVPR 2026] HiF-VLA: An efficient, bidirectional spatiotemporal expansion Vision-Language-Action Model☆75Mar 11, 2026Updated 4 months ago
- [ICLR 2026] Unified Vision-Language-Action Model☆316Oct 15, 2025Updated 9 months ago
- RynnVLA-002: A Unified Vision-Language-Action and World Model☆1,106Dec 2, 2025Updated 8 months ago
- Cosmos Policy☆847Jan 23, 2026Updated 6 months ago
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation☆59Mar 25, 2026Updated 4 months ago
- [RSS 2026] Causal video-action world model for generalist robot control☆1,718Jul 9, 2026Updated 3 weeks ago
- Self-CorrectingVLA:OnlineActionRefinementviaSparseWorldImagination☆26Apr 14, 2026Updated 3 months ago
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing☆3,374Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆43Jun 30, 2026Updated last month
- Official code of Motus: A Unified Latent Action World Model☆1,220Jan 5, 2026Updated 6 months ago
- The official implementation of MaskGRPO: Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models. (ICLR 2026, arxiv…☆19Jan 27, 2026Updated 6 months ago
- Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?☆1,237Apr 3, 2026Updated 4 months ago
- An all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.☆588Updated this week
- Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals☆2,516Apr 19, 2026Updated 3 months ago
- This Repository will be used by the Concordia Shanghai Swarm Team Club to share & work on different programs for the DJI TT Swarm Kit☆10Dec 19, 2021Updated 4 years ago
- [CVPR'26] QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models☆42Mar 25, 2026Updated 4 months ago
- [ICRA'25] Official code repository of "QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning"☆21Jun 25, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- LIBERO-PRO is the official repository of the LIBERO-PRO — an evaluation extension of the original LIBERO benchmark☆290Jul 17, 2026Updated 2 weeks ago
- InternVLA-A1: Unifying Understanding, Generation, and Action for Robotic Manipulation☆528Jul 20, 2026Updated 2 weeks ago
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success☆1,324Sep 9, 2025Updated 10 months ago
- Implementation of "SimVLA: A Simple VLA Baseline for Robotic Manipulation"☆141Feb 26, 2026Updated 5 months ago
- Official implementation of ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.☆269Apr 1, 2026Updated 4 months ago
- GeRM: A Generalist Robotic Model with Mixture-of-Experts for Quadruped Robot https://songwxuan.github.io/GeRM/☆37Apr 29, 2025Updated last year
- RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation☆75Jul 16, 2026Updated 2 weeks ago