[ACL 2024] GroundingGPT: Language-Enhanced Multi-modal Grounding Model
β342Nov 4, 2024Updated last year
Alternatives and similar repositories for GroundingGPT
Users that are interested in GroundingGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR 2024 π₯] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses thaβ¦β966Aug 5, 2025Updated last year
- [CVPR'2024 Highlight] Official PyTorch implementation of the paper "VTimeLLM: Empower LLM to Grasp Video Moments".β295Jun 13, 2024Updated 2 years ago
- UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Modelβ22Aug 5, 2024Updated 2 years ago
- The code of the paper "NExT-Chat: An LMM for Chat, Detection and Segmentation".β253Feb 5, 2024Updated 2 years ago
- [AAAI 2025] VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Groundingβ130Dec 10, 2024Updated last year
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- β404Jul 29, 2024Updated 2 years ago
- Official implementation of HawkEye: Training Video-Text LLMs for Grounding Text in Videosβ47Apr 29, 2024Updated 2 years ago
- [CVPR 2024] TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understandingβ425May 8, 2025Updated last year
- β134Dec 22, 2023Updated 2 years ago
- The official repository of "Video assistant towards large language model makes everything easy"β232Dec 24, 2024Updated last year
- β816Jul 8, 2024Updated 2 years ago
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)β861Jul 29, 2024Updated 2 years ago
- [ECCV 22] LocVTP: Video-Text Pre-training for Temporal Localizationβ39Jul 29, 2022Updated 4 years ago
- Official code for Goldfish model for long video understanding and MiniGPT4-video for short video understandingβ636Dec 10, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository contains the dataset, codebase, and benchmarks for our paper: <CNVid-3.5M: Build, Filter, and Pre-train the Large-scale Pβ¦β26Nov 28, 2023Updated 2 years ago
- [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of β¦β506Aug 9, 2024Updated 2 years ago
- [CVPR 2024] OneLLM: One Framework to Align All Modalities with Languageβ666Oct 22, 2024Updated last year
- [ACL 2024 π₯] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capβ¦β1,507Aug 5, 2025Updated last year
- Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Modelβ281Jun 25, 2024Updated 2 years ago
- [ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenizationβ585Jun 7, 2024Updated 2 years ago
- γEMNLP 2024π₯γVideo-LLaVA: Learning United Visual Representation by Alignment Before Projectionβ3,499Dec 3, 2024Updated last year
- PG-Video-LLaVA: Pixel Grounding in Large Multimodal Video Modelsβ264Aug 5, 2025Updated last year
- β4,711Jun 15, 2026Updated last month
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- γTMM 2025π₯γ Mixture-of-Experts for Large Vision-Language Modelsβ2,322Jul 15, 2025Updated last year
- [ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Textβ426May 5, 2025Updated last year
- The official code of Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval (AAAI2024)β32Mar 29, 2024Updated 2 years ago
- Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videosβ30Jun 24, 2024Updated 2 years ago
- [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understandingβ706Jan 29, 2025Updated last year
- γNeurIPS 2024γDense Connector for MLLMsβ182Oct 14, 2024Updated last year
- [CVPR 2024] Context-Guided Spatio-Temporal Video Groundingβ66Jun 28, 2024Updated 2 years ago
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactionsβ2,924May 26, 2025Updated last year
- The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025β279May 26, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β12Jul 23, 2024Updated 2 years ago
- Long Context Transfer from Language to Visionβ408Mar 18, 2025Updated last year
- [ECCV2024] Video Foundation Models & Data for Multimodal Understandingβ2,353Jul 2, 2026Updated last month
- β159Oct 31, 2024Updated last year
- [ICLR2026] VideoChat-Flash: Hierarchical Compression for Long-Context Video Modelingβ527Jul 19, 2026Updated 3 weeks ago
- Emu Series: Generative Multimodal Models from BAAIβ1,777Jan 12, 2026Updated 7 months ago
- [CVPR 2024] PixelLM is an effective and efficient LMM for pixel-level reasoning and understanding.β274Feb 11, 2025Updated last year