[CVPR 2026 (Highlight)] 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
☆40Jun 11, 2026Updated 3 months ago
Alternatives and similar repositories for 4D-RGPT
Users that are interested in 4D-RGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆12Mar 1, 2023Updated 3 years ago
- Official code of "Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer"☆20May 28, 2026Updated 4 months ago
- OGScene3D: Incremental Open-Vocabulary 3D Gaussian Scene Graph Mapping for Scene Understanding☆17Mar 18, 2026Updated 6 months ago
- (CVPR 2024) "Unsegment Anything by Simulating Deformation"☆29May 27, 2024Updated 2 years ago
- (3DV 2026 Oral) L4P -- a feed-forward foundational model designed for multiple low-level 4D vision perception tasks.☆76Dec 9, 2025Updated 9 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆76Jan 8, 2025Updated last year
- [RA-L'24, IROS'24] Official PyTorch Implementation of "Uni-DVPS: Unified Model for Depth-Aware Video Panoptic Segmentation"☆12Sep 20, 2026Updated last week
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data☆19Jun 2, 2026Updated 4 months ago
- [CVPR 2026] HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models☆42Jul 2, 2026Updated 3 months ago
- ☆13Mar 28, 2025Updated last year
- Pure C piecewise jerk path optimizer of Apollo for S-L Planning☆12Mar 26, 2024Updated 2 years ago
- Evaluation script for RoboSpatial-Home, a benchmark for spatial reasoning in 2D and 3D vision-language models.☆23May 14, 2026Updated 4 months ago
- Sa2VA-i is an improved version of the popular Sa2VA model☆17Nov 25, 2025Updated 10 months ago
- [NeurIPS 2024] Artemis: Towards Referential Understanding in Complex Videos☆27Apr 8, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [RAL 2026] GSO-SLAM☆27Jun 20, 2026Updated 3 months ago
- [CVPR 2026]"Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World"☆17Jul 7, 2026Updated 2 months ago
- [NeurIPS 2026] No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos☆98Sep 6, 2026Updated 3 weeks ago
- Prefect integrations with Microsoft Planetary Computer.☆10Jul 15, 2024Updated 2 years ago
- [ICCV 2025] V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction☆57Dec 2, 2025Updated 10 months ago
- Localization via embodied dialog on the navigation graph☆15Apr 18, 2022Updated 4 years ago
- [CVPR 2026] Official implementation of "GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation"☆23May 25, 2026Updated 4 months ago
- Package for the teleoperation of UR5+Allegro Robot composite☆12Apr 28, 2022Updated 4 years ago
- A Marketplace with Multi-Vendor based on Medusa that will be in Production ASAP☆10Apr 20, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICIP 2025 Best Student Paper Award] Official code release for "Pose-free 3D Gaussian splatting via shape-ray estimation"☆23Apr 2, 2026Updated 6 months ago
- CVPR 2025' Instruct-4DGS: Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic Separation☆36Sep 21, 2025Updated last year
- [ICCV2025] All in One: Visual-Description-Guided Unified Point Cloud Segmentation☆34Jul 25, 2025Updated last year
- Vehicle Trajectory Prediction Library☆16Feb 5, 2024Updated 2 years ago
- ☆29Apr 8, 2025Updated last year
- ☆74Apr 8, 2026Updated 5 months ago
- [NeurIPS 2025] SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models.☆18May 26, 2026Updated 4 months ago
- [AAAI 2026 Oral] STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes☆17Jan 23, 2026Updated 8 months ago
- ☆10Jun 26, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICCV 2025] VLM4D: Towards Spatiotemporal Awareness in Vision Language Models☆56Nov 20, 2025Updated 10 months ago
- Spatial Aptitude Training for Multimodal Langauge Models☆34Feb 8, 2026Updated 7 months ago
- [AAAI 2026 Oral] VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning☆27Mar 6, 2026Updated 6 months ago
- Using Kolmogorov Arnold Networks (KANs) instead of MLPs in PointNet for Classification and Segmentation of 3D Point Sets☆15Apr 23, 2026Updated 5 months ago
- Holistic Evaluation of Multimodal LLMs on Spatial Intelligence☆130Jul 1, 2026Updated 3 months ago
- [arXiv 2025] SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning☆70Dec 17, 2025Updated 9 months ago
- NX workspace for running medusa backend, storefront and admin panel with marketplace functionalities☆16Oct 6, 2022Updated 3 years ago