[TPAMI2025] Improving Generalized Visual Grounding with Instance-aware Joint Learning
☆33Apr 28, 2026Updated 4 months ago
Alternatives and similar repositories for InstanceVG
Users that are interested in InstanceVG are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICCV2025] PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination☆32Oct 13, 2025Updated 11 months ago
- [ECCV2026] MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding☆25Jun 19, 2026Updated 3 months ago
- [AAAI2025 selected as oral] - Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints☆45Jul 2, 2025Updated last year
- [PR2026] Drone Referring Localization: An Efficient Heterogeneous Spatial Feature Interaction Method For UAV Self-Localization☆97Feb 19, 2026Updated 7 months ago
- [NeurIPS2024] - SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion☆104Oct 29, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ECCV2024]FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance☆18Sep 11, 2024Updated 2 years ago
- 「TCSVT2021」A Transformer-Based Feature Segmentation and Region Alignment Method For UAV-View Geo-Localization☆122Mar 7, 2024Updated 2 years ago
- ☆13Oct 24, 2023Updated 2 years ago
- This is the implementation of the paper "Test-Time Adaptive Object Detection with Foundation Model" (Neurips 2025)☆23Jan 30, 2026Updated 7 months ago
- Code release for "Strike a Balance in Continual Panoptic Segmentation" (ECCV 2024)☆14Mar 14, 2025Updated last year
- Official Implementation of "OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation" (NeurIPS 2025).☆16Feb 27, 2026Updated 6 months ago
- Target-Grounded Graph-Aware Transformer for Aerial Vision-and-Dialog Navigation, AVDN Challenge, ICCV CLVL 2023.☆21Jan 2, 2024Updated 2 years ago
- paper list on Video Moment Retrieval (VMR), or Temporal Video Grounding (TVG), Video Grounding (VG), or Temporal Sentence Grounding in Vi…☆45Jul 30, 2026Updated last month
- EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing☆27May 29, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [WACV 2026] Official implementation of the paper: “CountingDINO: A Training-free Pipeline for Exemplar-based Class-Agnostic Counting”☆65Jun 22, 2026Updated 2 months ago
- [ICCV 2025] MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation☆23Sep 5, 2025Updated last year
- LLaVA-Next for STVG☆21Dec 5, 2025Updated 9 months ago
- RefDrone: A Challenging Benchmark for Drone Scene Referring Expression Comprehension☆48Jul 8, 2026Updated 2 months ago
- This is the open-sourced link of the TPAMI 2026 paper "SkyFind: A Large-Scale Benchmark Unveiling Referring Expression Comprehension for …☆38Aug 12, 2026Updated last month
- ☆28Feb 21, 2025Updated last year
- [TPAMI 2025] Towards Visual Grounding: A Survey☆326Nov 18, 2025Updated 10 months ago
- [CVPR2024] GSVA: Generalized Segmentation via Multimodal Large Language Models☆167Sep 12, 2024Updated 2 years ago
- Official repository for VCoT-Grasp.☆22Nov 18, 2025Updated 10 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision☆47Oct 19, 2025Updated 10 months ago
- FELA: Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation, AAAI 2025.☆38Dec 18, 2024Updated last year
- The first spoken long-text dataset derived from live streams, designed to reflect the redundancy-rich and conversational nature of real-w…☆12Jun 28, 2025Updated last year
- When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object Detection☆19Nov 3, 2025Updated 10 months ago
- ☆16May 9, 2024Updated 2 years ago
- An implementation of improved incremental Singular Value Decomposition(iSVD) algorithm☆17Jun 3, 2026Updated 3 months ago
- ☆29Apr 8, 2025Updated last year
- Improving One-stage Visual Grounding by Recursive Sub-query Construction, ECCV 2020☆90Sep 30, 2021Updated 4 years ago
- ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding☆20Aug 8, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Incrementally compute an approximate truncated singular value decomposition☆20Jun 29, 2026Updated 2 months ago
- ☆21Sep 16, 2025Updated last year
- 😎 up-to-date & curated list of awesome 3D Visual Grounding papers, methods & resources.☆284Jan 14, 2026Updated 8 months ago
- Pytorch implementation of Each Part Matters: Local Patterns Facilitate Cross-view Geo-localization https://arxiv.org/abs/2008.11646☆99Jul 6, 2026Updated 2 months ago
- Multimodal fusion☆22Dec 25, 2024Updated last year
- (CVPR25) Exploring Contextual Attribute Density in Referring Expression Counting☆20Dec 3, 2025Updated 9 months ago
- Official implementaiton of RefAM: Attention Magnets for Zero-Shot Referral Segmentaiton☆17Feb 6, 2026Updated 7 months ago