Sa2VA-i is an improved version of the popular Sa2VA model
☆16Nov 25, 2025Updated 7 months ago
Alternatives and similar repositories for Sa2VA-i
Users that are interested in Sa2VA-i are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [GCPR 2023] UGainS: Uncertainty Guided Anomaly Instance Segmentation☆16Jul 31, 2024Updated last year
- yolov3 ros node using tensorrt acceleration☆14Jan 27, 2023Updated 3 years ago
- [CVPR 2025] Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving☆46Nov 20, 2025Updated 8 months ago
- The website of the Objects365 Dataset☆13Jun 15, 2023Updated 3 years ago
- [ICLR 2025] Official Pytorch Implementation of MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segm…☆28Apr 3, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This is a repository contains the implementation of our NeurIPS'24 paper "Temporal Sentence Grounding with Relevance Feedback in Videos"☆13Aug 22, 2025Updated 10 months ago
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation (CVPR'25)☆21Apr 2, 2025Updated last year
- [ICLR2023] Video Scene Graph Generation from Single-Frame Weak Supervision☆12Sep 17, 2023Updated 2 years ago
- [CVPR 2024] Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations☆24Jan 20, 2025Updated last year
- Augmentation package for 3d data based on albumentaitons☆43Jun 18, 2025Updated last year
- Official code for Latest Object Memory Management for Temporally Consistent Video Instance Segmentation☆34Sep 17, 2025Updated 10 months ago
- PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation [NeurIPS 2025]☆19Oct 11, 2025Updated 9 months ago
- [ICCV 2023 Workshop] The Official Implementation of The First Prize Solution for RVOS Competition☆14Jan 1, 2024Updated 2 years ago
- [RA-L'24, IROS'24] Official PyTorch Implementation of "Uni-DVPS: Unified Model for Depth-Aware Video Panoptic Segmentation"☆13Oct 11, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Inferring and Leveraging Parts from Object Shape for Improving Semantic Image Synthesis (CVPR 2023)☆18Dec 13, 2024Updated last year
- ALGM applied to Segmenter☆33May 27, 2024Updated 2 years ago
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data☆19Jun 2, 2026Updated last month
- 《多模态大模型部署微调指南》快速部署/微调多模态大模型☆14Dec 4, 2024Updated last year
- [CVPR 2026]"Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World"☆17Jul 7, 2026Updated last week
- Online video temporal grounding☆16Oct 20, 2025Updated 9 months ago
- [ECCV24] VISA: Reasoning Video Object Segmentation via Large Language Model☆213Aug 5, 2024Updated last year
- Official code of Veason-R1☆15Jul 14, 2026Updated last week
- [NeurIPS 2025] SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models.☆17May 26, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Temporal Attention Fusion Network with Custom Loss Function for EEG-fNIRS Classification☆15Mar 18, 2026Updated 4 months ago
- Official repo for "DynaMITe: Dynamic Query Bootstrapping for Multi-object Interactive Segmentation Transformer"☆19Sep 29, 2023Updated 2 years ago
- LiVOS: Light Video Object Segmentation with Gated Linear Matching (CVPR 2025)☆48Sep 1, 2025Updated 10 months ago
- [CVPR 2026] Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation☆35Mar 12, 2026Updated 4 months ago
- The official fine-tuning code for DropTrack☆41Jul 1, 2023Updated 3 years ago
- 【干货】史上最全的PyTorch学习资源汇总☆11Aug 14, 2019Updated 6 years ago
- Learning Large-scale Neural Fields via Context Pruned Meta-Learning (NeurIPS 2023)☆28Sep 24, 2023Updated 2 years ago
- Use PYTHON to analyze data on Taobao user behavior. 利用python对淘宝用户行为进行数据分析。☆16May 21, 2020Updated 6 years ago
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models☆19Jan 21, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [CVPR 2025] Official PyTorch Implementation of GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmenta…☆70Jun 23, 2025Updated last year
- This is the official implementation of RGNet: A Unified Retrieval and Grounding Network for Long Videos☆20Mar 3, 2025Updated last year
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- Official Repo for CVPR 2025 Paper -- DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos☆17Mar 16, 2026Updated 4 months ago
- Mask4Former: Mask Transformer for 4D Panoptic Segmentation☆73May 28, 2025Updated last year
- ☆140Jul 4, 2024Updated 2 years ago
- VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.☆18May 31, 2026Updated last month