Official Implementation for paper "Referring Transformer: A One-step Approach to Multi-task Visual Grounding" Neurips 2021
☆67May 26, 2022Updated 4 years ago
Alternatives and similar repositories for RefTR
Users that are interested in RefTR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning, CVPR 2022☆97Dec 2, 2022Updated 3 years ago
- Encoder Fusion Network with Co-Attention Embedding for Referring Image Segmentation, CVPR2021☆21Aug 17, 2021Updated 4 years ago
- ☆234Apr 13, 2023Updated 3 years ago
- ☆198Feb 27, 2024Updated 2 years ago
- ☆41Jun 3, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- IROS 2023 "VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes"☆61Apr 22, 2024Updated 2 years ago
- SeqTR: A Simple yet Universal Network for Visual Grounding☆144Oct 30, 2024Updated last year
- Improving One-stage Visual Grounding by Recursive Sub-query Construction, ECCV 2020☆91Sep 30, 2021Updated 4 years ago
- ☆10Jan 9, 2025Updated last year
- [CVPR2020] Multi-task Collaborative Network for Joint Referring Expression Comprehension and Segmentation, CVPR2020 (oral)☆139Aug 4, 2022Updated 3 years ago
- awesome grounding: A curated list of research papers in visual grounding☆1,127Sep 21, 2025Updated 10 months ago
- [ICCV2021 & TPAMI2023] Vision-Language Transformer and Query Generation for Referring Segmentation☆366Jan 7, 2022Updated 4 years ago
- A curated list of research papers in Referring Expression Comprehension (REC)☆47May 13, 2021Updated 5 years ago
- A benchmark dataset for GREx: GRES, GREC, and GREG [CVPR 2023 & IJCV 2026]☆241Nov 14, 2025Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Lightweight Transformer for Multi-modal Tasks☆16Dec 9, 2022Updated 3 years ago
- A collection of papers about Referring Image Segmentation.☆826Jan 28, 2026Updated 5 months ago
- Referring Expression Parser☆27Feb 10, 2018Updated 8 years ago
- [CVPR2022] Official Implementation of ReferFormer☆355Feb 15, 2025Updated last year
- [CVPR-2023] The official dataset of Advancing Visual Grounding with Scene Knowledge: Benchmark and Method.☆34Jul 12, 2023Updated 3 years ago
- [ACM MM 22] Correspondence Matters for Video Referring Expression Comprehension☆15Sep 4, 2022Updated 3 years ago
- [CVPR2021] Look before you leap: learning landmark features for one-stage visual grounding.☆51Aug 31, 2021Updated 4 years ago
- ☆13Oct 30, 2023Updated 2 years ago
- ☆27Oct 7, 2021Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official codebase for "Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding"☆22Dec 20, 2020Updated 5 years ago
- ☆14Nov 4, 2022Updated 3 years ago
- [ICRA 2025] A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping☆13Feb 7, 2025Updated last year
- [AAAI 2023] DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding☆58Nov 28, 2022Updated 3 years ago
- Official code for the paper, "TaCA: Upgrading Your Visual Foundation Model with Task-agnostic Compatible Adapter".☆16Jun 20, 2023Updated 3 years ago
- ☆16Nov 14, 2022Updated 3 years ago
- ☆1,050Oct 3, 2022Updated 3 years ago
- An official PyTorch implementation of the CRIS paper☆281Jun 9, 2024Updated 2 years ago
- ☆47Oct 3, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Source code for EMNLP 2022 paper “PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models”☆49Nov 10, 2022Updated 3 years ago
- Code for Learned Thresholds Token Merging and Pruning for Vision Transformers (LTMP). A technique to reduce the size of Vision Transforme…☆17Nov 24, 2024Updated last year
- This is an implementation of "Grounding of Textual Phrases in Images by Reconstruction" in PyTorch☆18Apr 7, 2020Updated 6 years ago
- ☆22Jun 20, 2024Updated 2 years ago
- iterative shrinking for referring expression grounding using deep reinforcement learning☆14Nov 27, 2021Updated 4 years ago
- Dataset API for "PhraseCut: Language-based Image Segmentation in the Wild"☆116Mar 28, 2026Updated 3 months ago
- Code for the paper titled "CiT Curation in Training for Effective Vision-Language Data".☆78Jan 18, 2023Updated 3 years ago