[ACM MM'26] MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
☆65May 14, 2026Updated 3 months ago
Alternatives and similar repositories for MMaDA-VLA
Users that are interested in MMaDA-VLA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'25] SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning☆40Oct 14, 2025Updated 10 months ago
- ActionCodec: What Makes for Good Action Tokenizers☆62Mar 1, 2026Updated 5 months ago
- [ICLR 2026] Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining☆32Apr 26, 2026Updated 3 months ago
- Official implementation of FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment☆55Mar 24, 2026Updated 5 months ago
- FASTER: Rethinking Real-Time Flow VLAs☆147May 14, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [ICML 2026] ResVLA: From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges☆29Jun 1, 2026Updated 2 months ago
- LAP: Language-Action Pre-Training Enables Zero-Shot Cross Embodiment Transfer☆165May 20, 2026Updated 3 months ago
- Dream-VL and Dream-VLA, a diffusion VLM and a diffusion VLA.☆114Jan 14, 2026Updated 7 months ago
- [RSS 2026] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion☆305May 26, 2026Updated 2 months ago
- 🔥 The first open-sourced diffusion vision-langauge-action model. [ICLR 2026]☆186Mar 12, 2026Updated 5 months ago
- Official implementation of Spatial-Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model [ICLR2026]☆281Jul 7, 2026Updated last month
- Reshaping Action Error Distributions for Reliable Vision-Language-Action Models☆17Feb 5, 2026Updated 6 months ago
- Code for Xiaomi-Robotics-1☆641Aug 17, 2026Updated last week
- ☆49May 12, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆33Jun 7, 2026Updated 2 months ago
- [CVPR 2026] HiF-VLA: An efficient, bidirectional spatiotemporal expansion Vision-Language-Action Model☆76Mar 11, 2026Updated 5 months ago
- [ICLR 2026] Unified Vision-Language-Action Model☆322Oct 15, 2025Updated 10 months ago
- RynnVLA-002: A Unified Vision-Language-Action and World Model☆1,117Dec 2, 2025Updated 8 months ago
- Cosmos Policy☆853Jan 23, 2026Updated 7 months ago
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation☆59Aug 10, 2026Updated last week
- [RSS 2026] Causal video-action world model for generalist robot control☆1,797Jul 9, 2026Updated last month
- Self-CorrectingVLA:OnlineActionRefinementviaSparseWorldImagination☆27Apr 14, 2026Updated 4 months ago
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing☆3,510Aug 9, 2026Updated 2 weeks ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆47Jun 30, 2026Updated last month
- Official code of Motus: A Unified Latent Action World Model☆1,238Jan 5, 2026Updated 7 months ago
- The official implementation of MaskGRPO: Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models. (ICLR 2026, arxiv…☆19Jan 27, 2026Updated 6 months ago
- Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?☆1,338Updated this week
- An all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.☆627Updated this week
- Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals☆2,578Apr 19, 2026Updated 4 months ago
- This Repository will be used by the Concordia Shanghai Swarm Team Club to share & work on different programs for the DJI TT Swarm Kit☆10Dec 19, 2021Updated 4 years ago
- [CVPR'26] QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models☆44Mar 25, 2026Updated 4 months ago
- [ICRA'25] Official code repository of "QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning"☆22Jun 25, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LIBERO-PRO is the official repository of the LIBERO-PRO — an evaluation extension of the original LIBERO benchmark☆301Jul 17, 2026Updated last month
- InternVLA-A1: Unifying Understanding, Generation, and Action for Robotic Manipulation☆539Jul 20, 2026Updated last month
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success☆1,343Sep 9, 2025Updated 11 months ago
- Implementation of "SimVLA: A Simple VLA Baseline for Robotic Manipulation"☆142Feb 26, 2026Updated 5 months ago
- Official implementation of ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.☆274Apr 1, 2026Updated 4 months ago
- GeRM: A Generalist Robotic Model with Mixture-of-Experts for Quadruped Robot https://songwxuan.github.io/GeRM/☆38Apr 29, 2025Updated last year
- RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation☆85Jul 16, 2026Updated last month