[ICLR 2025] GRAM Official PyTorch repository
☆138May 7, 2025Updated last year
Alternatives and similar repositories for GRAM
Users that are interested in GRAM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] Multi-modal representation learning of shared, unique and synergistic features between modalities☆72May 6, 2025Updated last year
- Official PyTorch repository for StawGAN: Structural-Aware Generative Adversarial Networks for Infrared Image Translation☆19Oct 18, 2023Updated 2 years ago
- Official PyTorch repository for Ship in Sight: Diffusion Models for Ship-Image Super Resolution, WCCI 2024.☆28Jan 7, 2025Updated last year
- [NeurIPS 2025 Spotlight] Implementation of the paper "InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interacti…☆28Jan 4, 2026Updated 7 months ago
- Official PyTorch repository for ICASSP 2025 Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion Models.☆73Aug 1, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Diffusion Models for Audio Semantic Communication☆18Apr 17, 2024Updated 2 years ago
- [ECCV 2024 Oral] Official implementation of the paper "DEVIAS: Learning Disentangled Video Representations of Action and Scene"☆29Nov 15, 2025Updated 9 months ago
- [ICLR 2025] This repo is the official implementation of our paper "Learning Fine-Grained Representations through Textual Token Disentangl…☆23Jul 28, 2025Updated last year
- [NIPS2023] Code and Model for VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset☆302Mar 14, 2024Updated 2 years ago
- ☆26Feb 29, 2024Updated 2 years ago
- An official implementation of "DiffPS: Leveraging Prior Knowledge of Diffusion Model for Person Search" (ICCV 2025 Highlight).☆17Mar 20, 2026Updated 5 months ago
- Official repository for "Boosting Audio Visual Question Answering via Key Semantic-Aware Cues" in ACM MM 2024.☆17Oct 25, 2024Updated last year
- ICML-2024 highlight paper "Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization"☆20Jul 18, 2024Updated 2 years ago
- Cross Modality Optimal Transport for multimodal inference☆10Nov 29, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆17Dec 4, 2024Updated last year
- The official code for Improving Multimodal Learning via Imbalanced Learning☆41Mar 26, 2026Updated 5 months ago
- Improving Medical Vision-Language Contrastive Pretraining with Semantics-aware Triage☆12Jun 25, 2023Updated 3 years ago
- [NeurIPS 2023] Factorized Contrastive Learning: Going Beyond Multi-view Redundancy☆78Nov 13, 2023Updated 2 years ago
- [2025 CVPR] Towards Open-Vocabulary Audio-Visual Event Localization☆47Mar 7, 2025Updated last year
- ☆43Mar 28, 2024Updated 2 years ago
- [CVPR 2025] 🔥 Official impl. of "Audio-Visual Instance Segmentation".☆52Jun 5, 2025Updated last year
- body25 + hand pose3d to bvh☆17Jul 19, 2022Updated 4 years ago
- ☆12Sep 15, 2024Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆50Nov 24, 2024Updated last year
- Vision Relation Transformer for Unbiased Scene Graph Generation (ICCV 2023)☆22Mar 23, 2026Updated 5 months ago
- LinVT: Empower Your Image-level Large Language Model to Understand Videos☆83Dec 30, 2024Updated last year
- [ICLR 2026] DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning☆141Jul 2, 2026Updated 2 months ago
- SFI-Swin: Symmetric Face Inpainting with Swin Transformer by Distinctly Learning Face Components Distributions https://arxiv.org/abs/23…☆15Jan 12, 2023Updated 3 years ago
- [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection☆30Sep 26, 2024Updated last year
- ☆21May 4, 2023Updated 3 years ago
- This repository contains the official implementation, data generation tools, and benchmark datasets for our research on synthetic data fo…☆16Apr 1, 2026Updated 5 months ago
- [ECCV'24 Oral] PiTe: Pixel-Temporal Alignment for Large Video-Language Model☆17Feb 13, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is the repo for "Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition", CVPR2025.☆27Dec 22, 2025Updated 8 months ago
- UMT is a unified and flexible framework which can handle different input modality combinations, and output video moment retrieval and/or …☆238Apr 15, 2024Updated 2 years ago
- Code and data release for the paper "Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Align…☆19Apr 5, 2024Updated 2 years ago
- ☆19Nov 11, 2025Updated 9 months ago
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference☆25May 21, 2026Updated 3 months ago
- Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval -- AAAI2025☆22May 8, 2026Updated 3 months ago
- [AAAI 2026 Oral] The official code of "UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning"☆75Dec 8, 2025Updated 8 months ago