Pytorch code for paper From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
β212Jan 8, 2025Updated last year
Alternatives and similar repositories for COMM
Users that are interested in COMM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β19Dec 6, 2023Updated 2 years ago
- [CVPR 2024 π₯] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses thaβ¦β968Sep 5, 2026Updated 2 weeks ago
- This repo contains the code for our paper Towards Open-Ended Visual Recognition with Large Language Modelβ102Jul 15, 2024Updated 2 years ago
- [TMM 2023] Self-paced Curriculum Adapting of CLIP for Visual Grounding.β135Nov 10, 2025Updated 10 months ago
- Official Implementation of ICCV 2023 Paper - SegPrompt: Boosting Open-World Segmentation via Category-level Prompt Learningβ110May 28, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [CVPR 2024] CapsFusion: Rethinking Image-Text Data at Scaleβ216Feb 27, 2024Updated 2 years ago
- β134Dec 22, 2023Updated 2 years ago
- Official implementation for the paper "Prompt Pre-Training with Over Twenty-Thousand Classes for Open-Vocabulary Visual Recognition"β259May 3, 2024Updated 2 years ago
- Official implementation of TagAlignβ37Dec 11, 2024Updated last year
- VisionLLM Seriesβ1,154Feb 27, 2025Updated last year
- Official implementation of 'CLIP-DINOiser: Teaching CLIP a few DINO tricks' paper.β285Oct 26, 2024Updated last year
- β59Aug 7, 2023Updated 3 years ago
- [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of β¦β506Aug 9, 2024Updated 2 years ago
- Code/Data for the paper: "LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding"β268Jun 12, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- β90Nov 25, 2023Updated 2 years ago
- Recognize Any Regionsβ123Dec 18, 2024Updated last year
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perceptionβ161Dec 6, 2024Updated last year
- β815Jul 8, 2024Updated 2 years ago
- Emu Series: Generative Multimodal Models from BAAIβ1,778Jan 12, 2026Updated 8 months ago
- [ICCV 2023] Official implementation of the paper "A Simple Framework for Open-Vocabulary Segmentation and Detection"β764Jan 22, 2024Updated 2 years ago
- [CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Promptsβ338Jul 17, 2024Updated 2 years ago
- EVA Series: Visual Representation Fantasies from BAAIβ2,694Aug 1, 2024Updated 2 years ago
- NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024β1,854Aug 11, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [CVPR 2024] Official implementation of the paper "Visual In-context Learning"β544Apr 8, 2024Updated 2 years ago
- A collection of visual instruction tuning datasets.β76Mar 14, 2024Updated 2 years ago
- Project Page for "LISA: Reasoning Segmentation via Large Language Model"β2,683Feb 16, 2025Updated last year
- [ICCV2023] VLPart: Going Denser with Open-Vocabulary Part Segmentationβ395Sep 19, 2023Updated 3 years ago
- [NeurIPS2023] Code release for "Hierarchical Open-vocabulary Universal Image Segmentation"β293Jun 19, 2025Updated last year
- [NeurIPS 2024] Official implementation of the paper "Interfacing Foundation Models' Embeddings"β133Aug 21, 2024Updated 2 years ago
- Grounded Language-Image Pre-trainingβ2,612Jan 24, 2024Updated 2 years ago
- Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMsβ100Jan 16, 2025Updated last year
- β129Jul 29, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official implementation of paper "MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens"β869May 8, 2025Updated last year
- (ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interestβ555Jun 3, 2025Updated last year
- [ECCV 2024] Tokenize Anything via Promptingβ600Dec 11, 2024Updated last year
- β91Jul 4, 2024Updated 2 years ago
- [CVPR 2024] Alpha-CLIP: A CLIP Model Focusing on Wherever You Wantβ875Jul 20, 2025Updated last year
- [CVPR 2023] Official Implementation of X-Decoder for generalized decoding for pixel, image and languageβ1,343Oct 5, 2023Updated 2 years ago
- [NeurIPS2023] Official implementation of the paper "Large Language Models are Visual Reasoning Coordinators"β106Nov 9, 2023Updated 2 years ago