[ICLR 2025] GRAM Official PyTorch repository
☆137May 7, 2025Updated last year
Alternatives and similar repositories for GRAM
Users that are interested in GRAM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] Multi-modal representation learning of shared, unique and synergistic features between modalities☆70May 6, 2025Updated last year
- Official PyTorch repository for Ship in Sight: Diffusion Models for Ship-Image Super Resolution, WCCI 2024.☆26Jan 7, 2025Updated last year
- Code & Weights for “Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation”☆15Dec 6, 2024Updated last year
- [NeurIPS 2025 Spotlight] Implementation of the paper "InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interacti…☆27Jan 4, 2026Updated 6 months ago
- [TPAMI 2026] Principled Multimodal Representation Learning☆37May 10, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ECCV 2024 Oral] Official implementation of the paper "DEVIAS: Learning Disentangled Video Representations of Action and Scene"☆29Nov 15, 2025Updated 8 months ago
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval (ICCV 2025 Highlight)☆26Aug 1, 2025Updated 11 months ago
- ☆22Jan 17, 2025Updated last year
- [ICLR 2025] This repo is the official implementation of our paper "Learning Fine-Grained Representations through Textual Token Disentangl…☆23Jul 28, 2025Updated 11 months ago
- This repository is an official implementation of AVIGATE (CVPR 2025, oral)☆19Aug 21, 2025Updated 11 months ago
- [NIPS2023] Code and Model for VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset☆302Mar 14, 2024Updated 2 years ago
- ICML-2024 highlight paper "Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization"☆19Jul 18, 2024Updated 2 years ago
- Official GitHub repository of the lecture "Multimodal Deep Learning for Recommendation", at the 2024 ACM RecSys Summer School☆12Oct 12, 2024Updated last year
- Cross Modality Optimal Transport for multimodal inference☆10Nov 29, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆17Dec 4, 2024Updated last year
- Official Repository for ICML 2024 Paper "OT-CLIP: Understanding and Generalizing CLIP via Optimal Transport"☆23Dec 4, 2025Updated 7 months ago
- Improving Medical Vision-Language Contrastive Pretraining with Semantics-aware Triage☆11Jun 25, 2023Updated 3 years ago
- [NeurIPS 2023] Factorized Contrastive Learning: Going Beyond Multi-view Redundancy☆76Nov 13, 2023Updated 2 years ago
- [2025 CVPR] Towards Open-Vocabulary Audio-Visual Event Localization☆46Mar 7, 2025Updated last year
- [WACV 2024] Code release for "VEATIC: Video-based Emotion and Affect Tracking in Context Dataset"☆24Jan 14, 2026Updated 6 months ago
- [CVPR 2025] 🔥 Official impl. of "Audio-Visual Instance Segmentation".☆52Jun 5, 2025Updated last year
- body25 + hand pose3d to bvh☆17Jul 19, 2022Updated 4 years ago
- Vision Relation Transformer for Unbiased Scene Graph Generation (ICCV 2023)☆22Mar 23, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection☆30Sep 26, 2024Updated last year
- ☆21May 4, 2023Updated 3 years ago
- [CVPR 2025] Enhanced OoD Detection through Cross-Modal Alignment of Multi-modal Representations☆32Jun 27, 2025Updated last year
- [ECCV'24 Oral] PiTe: Pixel-Temporal Alignment for Large Video-Language Model☆17Feb 13, 2025Updated last year
- final-project-level3-nlp-02 created by GitHub Classroom☆11Dec 31, 2021Updated 4 years ago
- Generating Human Skeletons with Mutual Actions☆11Oct 22, 2021Updated 4 years ago
- This is the repo for "Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition", CVPR2025.☆24Dec 22, 2025Updated 7 months ago
- UMT is a unified and flexible framework which can handle different input modality combinations, and output video moment retrieval and/or …☆238Apr 15, 2024Updated 2 years ago
- A collection of paper or resources for drag editing☆22Apr 23, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code and data release for the paper "Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Align…☆19Apr 5, 2024Updated 2 years ago
- Official code release of "DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding" [ICCV2025 Highlight]☆52Sep 27, 2025Updated 9 months ago
- ☆21Nov 11, 2025Updated 8 months ago
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference☆25May 21, 2026Updated 2 months ago
- Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval -- AAAI2025☆21May 8, 2026Updated 2 months ago
- [AAAI 2026 Oral] The official code of "UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning"☆74Dec 8, 2025Updated 7 months ago
- ☆14Apr 9, 2026Updated 3 months ago