[ICLR 2025] GRAM Official PyTorch repository
☆138May 7, 2025Updated last year
Alternatives and similar repositories for GRAM
Users that are interested in GRAM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Apr 24, 2026Updated 3 months ago
- Code & Weights for “Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation”☆15Dec 6, 2024Updated last year
- [NeurIPS 2025 Spotlight] Implementation of the paper "InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interacti…☆27Jan 4, 2026Updated 7 months ago
- [ECCV 2024 Oral] Official implementation of the paper "DEVIAS: Learning Disentangled Video Representations of Action and Scene"☆29Nov 15, 2025Updated 8 months ago
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval (ICCV 2025 Highlight)☆27Aug 1, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆22Jan 17, 2025Updated last year
- [ICLR 2025] This repo is the official implementation of our paper "Learning Fine-Grained Representations through Textual Token Disentangl…☆23Jul 28, 2025Updated last year
- This repository is an official implementation of AVIGATE (CVPR 2025, oral)☆19Aug 21, 2025Updated 11 months ago
- An official implementation of "DiffPS: Leveraging Prior Knowledge of Diffusion Model for Person Search" (ICCV 2025 Highlight).☆17Mar 20, 2026Updated 4 months ago
- Official repository for "Boosting Audio Visual Question Answering via Key Semantic-Aware Cues" in ACM MM 2024.☆17Oct 25, 2024Updated last year
- ICML-2024 highlight paper "Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization"☆19Jul 18, 2024Updated 2 years ago
- Official GitHub repository of the lecture "Multimodal Deep Learning for Recommendation", at the 2024 ACM RecSys Summer School☆12Oct 12, 2024Updated last year
- Cross Modality Optimal Transport for multimodal inference☆10Nov 29, 2024Updated last year
- ☆17Dec 4, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official Repository for ICML 2024 Paper "OT-CLIP: Understanding and Generalizing CLIP via Optimal Transport"☆23Dec 4, 2025Updated 8 months ago
- The official code for Improving Multimodal Learning via Imbalanced Learning☆42Mar 26, 2026Updated 4 months ago
- [2025 CVPR] Towards Open-Vocabulary Audio-Visual Event Localization☆46Mar 7, 2025Updated last year
- ☆43Mar 28, 2024Updated 2 years ago
- [CVPR 2025] 🔥 Official impl. of "Audio-Visual Instance Segmentation".☆52Jun 5, 2025Updated last year
- Vision Relation Transformer for Unbiased Scene Graph Generation (ICCV 2023)☆22Mar 23, 2026Updated 4 months ago
- LinVT: Empower Your Image-level Large Language Model to Understand Videos☆83Dec 30, 2024Updated last year
- [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection☆30Sep 26, 2024Updated last year
- ☆21May 4, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2025] Enhanced OoD Detection through Cross-Modal Alignment of Multi-modal Representations☆32Jun 27, 2025Updated last year
- This repository contains the official implementation, data generation tools, and benchmark datasets for our research on synthetic data fo…☆15Apr 1, 2026Updated 4 months ago
- This is the repo for "Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition", CVPR2025.☆26Dec 22, 2025Updated 7 months ago
- UMT is a unified and flexible framework which can handle different input modality combinations, and output video moment retrieval and/or …☆238Apr 15, 2024Updated 2 years ago
- Code and data release for the paper "Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Align…☆19Apr 5, 2024Updated 2 years ago
- Missing Modality Generation for Recommendaton☆37Aug 1, 2026Updated last week
- ☆20Nov 11, 2025Updated 9 months ago
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference☆25May 21, 2026Updated 2 months ago
- Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval -- AAAI2025☆21May 8, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [AAAI 2026 Oral] The official code of "UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning"☆74Dec 8, 2025Updated 8 months ago
- ☆14Apr 9, 2026Updated 4 months ago
- ☆29Jul 25, 2025Updated last year
- ☆19Jul 28, 2025Updated last year
- [PR 2024] A large Cross-Modal Video Retrieval Dataset with Reading Comprehension☆32Dec 28, 2023Updated 2 years ago
- ☆23Dec 1, 2025Updated 8 months ago
- Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representation Learning for Temporal Gr…☆154Aug 21, 2024Updated last year