A curated list of academic papers and resources on Vision-Language-Action (VLA) and World Action Models (WAM)
☆29Aug 3, 2026Updated this week
Alternatives and similar repositories for Awesome-World-Action-Model
Users that are interested in Awesome-World-Action-Model are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation☆67Jul 28, 2026Updated last week
- [ICCV 23] A Simple Vision Transformer for Weakly Semi-supervised 3D Object Detection☆13Apr 12, 2024Updated 2 years ago
- [CVPR 2026] PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding☆35Apr 7, 2026Updated 3 months ago
- [CVPR 2026] When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models☆68Apr 11, 2026Updated 3 months ago
- [ECCV 2024] Make Your ViT-based Multi-view 3D Detectors Faster via Token Compression☆53Sep 21, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ECCV 2026] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding☆421Jun 18, 2026Updated last month
- 钢材表面缺陷检测与分割竞赛的解决方案☆23Nov 12, 2024Updated last year
- Next Forcing: World Action Modeling with Multi-Chunk Prediction (MCP)☆114Jul 19, 2026Updated 2 weeks ago
- the official code of DriveMonkey☆45Mar 20, 2026Updated 4 months ago
- Awesome GPT-4 with Applications. This is a collection of resources related to GPT-4, including news, official documents, demo and applica…☆20Mar 15, 2023Updated 3 years ago
- Code for "Reconstructing 3D Human Pose from RGB-D Data with Occlusions" (PG 2023)☆13Nov 5, 2023Updated 2 years ago
- Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?☆1,237Apr 3, 2026Updated 4 months ago
- ☆43Jun 30, 2026Updated last month
- URDF-based forward and inverse kinematics helpers for robot arms — pluggable IK backends (Placo, PyRoki, RoboPlan), Viser 3D visualizatio…☆23May 22, 2026Updated 2 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICLR26] ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding☆95Mar 20, 2026Updated 4 months ago
- Isaac Lab implementation of AMP(Adversarial Motion Prior) with rl_games☆14Aug 5, 2025Updated 11 months ago
- Live 3D viewer and editor for MuJoCo MJCF models, inside VS Code☆27Jul 22, 2026Updated last week
- [CVPR 2025] A Unified Image-Dense Annotation Generation Model for Underwater Scenes☆60Apr 9, 2025Updated last year
- Official code repository of Shuffle-R1☆26Feb 23, 2026Updated 5 months ago
- Vega: Learning to Drive with Natural Language Instructions☆43Mar 27, 2026Updated 4 months ago
- A light-weight, Eigen-based C++ library for trajectory optimization for legged robots.☆27Aug 11, 2021Updated 4 years ago
- Localized Vision-Language Matching for Open-vocabulary Object Detection☆22Aug 11, 2022Updated 3 years ago
- ☆17Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ECCV 26] Video Streaming Thinking☆120Jul 28, 2026Updated last week
- Code for ICLR 2022 publication: Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RL. https://openreview…☆10Aug 31, 2024Updated last year
- 华中科技大学人工智能与自动化学院课程资源☆67May 21, 2023Updated 3 years ago
- Replication of mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs☆27Apr 13, 2026Updated 3 months ago
- This is a PyTorch implementation of MCLN proposed by our paper "Multi-branch Collaborative Learning Network for 3D Visual Grounding"(ECCV…☆27Oct 10, 2024Updated last year
- ☆32Jul 10, 2026Updated 3 weeks ago
- [ICML26] Official Repo for WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching☆40Jul 23, 2026Updated last week
- [NeurIPS 2024] A Unified Framework for 3D Scene Understanding☆179Jul 7, 2025Updated last year
- [ICLR26] Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning☆21Apr 16, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of HEAD CoRL 2025☆27Aug 22, 2025Updated 11 months ago
- [MM2024 Oral] 3D-GRES: Generalized 3D Referring Expression Segmentation☆43Dec 15, 2024Updated last year
- [CVPR 2024] Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud Analysis☆172Oct 11, 2024Updated last year
- ☆10Jul 25, 2016Updated 10 years ago
- Dataset for EMNLP'23 Paper "DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine Reading"☆11Oct 25, 2023Updated 2 years ago
- ☆30Jun 29, 2026Updated last month
- [ICCV 2025] HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation☆259May 12, 2026Updated 2 months ago