[TMM 2025] This is the official Pytorch code for our paper "Visual Position Prompt for MLLM based Visual Grounding".
☆32Jul 23, 2025Updated last year
Alternatives and similar repositories for VPP-LLaVA
Users that are interested in VPP-LLaVA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Oct 13, 2025Updated 10 months ago
- [TPAMI 2024] This is the official Pytorch code for our paper "Context Disentangling and Prototype Inheriting for Robust Visual Grounding"…☆28May 8, 2025Updated last year
- 16k Hz Vocoder (HiFiGAN Codes and Pretrained Models)☆18Apr 3, 2023Updated 3 years ago
- [Symmetry 2019] This is the Matlab code for our paper "Optimizing MSE for Clustering with Balanced Size Constraints".☆20Mar 25, 2025Updated last year
- ☆17Jul 14, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2025] This repo is the official implementation of our paper "Learning Fine-Grained Representations through Textual Token Disentangl…☆23Jul 28, 2025Updated last year
- The code for AAAI 2025 “Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation”☆15Jan 3, 2025Updated last year
- Vision-Language based Visual Object Tracking☆36Jul 4, 2026Updated 2 months ago
- ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memory☆67Apr 21, 2026Updated 4 months ago
- ☆16Apr 4, 2022Updated 4 years ago
- A C++ implementation of stft, melspectrogram and mel_to_stft☆11Jun 2, 2022Updated 4 years ago
- ☆13May 27, 2026Updated 3 months ago
- Implementation of our paper "Exploiting Unsupervised Data for Emotion Recognition in Conversations" in the Findings of EMNLP-2020.☆13Nov 17, 2020Updated 5 years ago
- [TIP] Exploring Effective Factors for Improving Visual In-Context Learning☆21Jul 2, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A modern macOS Pomodoro & focus tool for deep focus.☆24Jul 13, 2026Updated last month
- Pytorch Implementation of Automatic Relation-aware Graph Network Proliferation (CVPR'22, Oral)☆26Apr 3, 2024Updated 2 years ago
- Segment Anything Model 2 C++ ONNX Wrapper☆18May 13, 2026Updated 3 months ago
- Waterbody style transfer of underwater imagery (JOE 2025)☆26Dec 12, 2025Updated 8 months ago
- [CVPR 2024] Code for "Improved Visual Grounding through Self-Consistent Explanations".☆28Mar 1, 2024Updated 2 years ago
- Project page for the 'CLAWS: Clustering Assisted Weakly Supervised Learning with Normalcy Suppression for Anomalous Event Detection', ECC…☆12May 29, 2021Updated 5 years ago
- [EMNLP 2024 Main] MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension☆16Jan 6, 2025Updated last year
- [ICCV'25] Official PyTorch Implementation of "VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models"☆17Dec 8, 2025Updated 8 months ago
- Know What and Know Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language Navigation☆16Feb 7, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation example of Distributed Tensorflow☆10Jul 22, 2017Updated 9 years ago
- [NeurIPS 2024] official code release for our paper "Revisiting the Integration of Convolution and Attention for Vision Backbone".☆43Jan 21, 2025Updated last year
- [CVPR'25] Attention IoU: Examining Biases in CelebA using Attention Maps☆13Mar 26, 2025Updated last year
- [TPAMI]CTNet: Context-based Tandem Network for Semantic Segmentation☆16Jun 15, 2022Updated 4 years ago
- ☆13May 21, 2024Updated 2 years ago
- [EMNLP'26] Code and data for VTCBench, a VLM benchmark for long-context understanding capabilities under vision-text compression paradigm…☆27Aug 27, 2026Updated last week
- Extract MFCCs from videos and make bag-of-audio-words (BOAW) representations.☆11Dec 20, 2018Updated 7 years ago
- ☆47Apr 16, 2026Updated 4 months ago
- [IEEE TMM] Code for the paper "HRNeXt: High-Resolution Context Network for Crowd Pose Estimation"☆10Feb 24, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Video Depth Propagation [3DV 2026]☆38Jan 23, 2026Updated 7 months ago
- Unofficial version of LaneExtraction☆14Oct 12, 2022Updated 3 years ago
- [ACL 2021] This is the Pytorch code for our paper "Semantic Relation-aware Difference Representation Learning for Change Captioning".☆13Jan 16, 2022Updated 4 years ago
- 计算机视觉☆13Nov 13, 2023Updated 2 years ago
- LITE: A Paradigm Shift: Multi Object Tracking with Deep Association Metric☆29Jun 15, 2026Updated 2 months ago
- Contains the implementation of the EDAIN and EDAIN-KL methods proposed in our paper. The research was also part of the thesis I wrote as …☆16Feb 19, 2024Updated 2 years ago
- Hierarchical Group Sparse Regularization for Deep Convolutional Neural Networks☆11Apr 13, 2020Updated 6 years ago