The implementation of FINER-MLLM, which is accepted by MM2024.
☆18Oct 8, 2024Updated last year
Alternatives and similar repositories for FINER-MLLM
Users that are interested in FINER-MLLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ECCV 2024] Official repository of "GenView: Enhancing View Quality with Pretrained Generative Model for Self-Supervised Learning".☆29Dec 18, 2024Updated last year
- Official repository of the "Fine-grained Key-Value Memory Enhanced Predictor for Video Representation Learning" (ACM MM 2023)☆23Jul 11, 2024Updated 2 years ago
- [ACM MM 2026] Official implementation of “Continuous Knowledge-Preserving Decomposition with Adaptive Layer Selection for Few-Shot Class-…☆34Jul 12, 2026Updated last month
- ICME2022 Special Session “Beyond Accuracy: Responsible, Responsive, and Robust Multimedia Retrieval ”☆12Jun 3, 2024Updated 2 years ago
- ☆31Jun 22, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Viewpoint-aware Attentive Multi-view Inference for Vehicle Re-identification☆13Mar 27, 2019Updated 7 years ago
- Official Implementation of CAPEAM (ICCV'23)☆16Nov 30, 2024Updated last year
- Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question Answering☆11Feb 16, 2023Updated 3 years ago
- Code of ACM MM 2023 Paper: A Symbolic Characters Aware Model for Solving Geometry Problems☆16Dec 27, 2023Updated 2 years ago
- This is the official code for "Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning"☆26Apr 2, 2026Updated 5 months ago
- A lightweight open-source package to fine-tune embedding models.☆22Feb 4, 2024Updated 2 years ago
- Add Rain Streak Mask On Unparied Image Using GAN☆10Sep 12, 2020Updated 5 years ago
- Repo for the paper "Words or Vision: Do Vision-Language Models Have Blind Faith in Text?" (CVPR 2025)☆18Mar 31, 2026Updated 5 months ago
- UAVM @ ACM MM2023 Workshop on UAVs in Multimedia: Capturing the World from a New Perspective☆17Apr 30, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Python library of cover tree (http://hunch.net/~jl/projects/cover_tree/cover_tree.html) for fast nearest neighbor querying☆16Jan 3, 2012Updated 14 years ago
- A python package of robust and effective defogging/dehazing method☆15Dec 30, 2018Updated 7 years ago
- [NeurIPS2023] LoRA: A Logical Reasoning Augmented Dataset for Visual Question Answering☆12Jan 5, 2024Updated 2 years ago
- ☆14Sep 28, 2024Updated last year
- ☆13Aug 14, 2022Updated 4 years ago
- ☆20Nov 4, 2023Updated 2 years ago
- Code and dataset for NAACL 2022 paper "CoSIm: Commonsense Reasoning for Counterfactual Scene Imagination" Hyounghun Kim, Abhay Zala, Mohi…☆16Nov 26, 2022Updated 3 years ago
- [ICLR 2023] This is the code repo for our ICLR‘23 paper "Universal Vision-Language Dense Retrieval: Learning A Unified Representation Spa…☆52Jul 3, 2024Updated 2 years ago
- Codes of the Fine-grained Textual Inversion network for Zero-Shot Composed Image Retrieval☆27Apr 9, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Python (pip) package for fitting mixtures of Student's t-distributions using either maximum likelihood (EM) or Bayesian methodology (vari…☆11Sep 23, 2025Updated 11 months ago
- A curated list of Story Ending Generation models; DASFAA'22: Incorporating Commonsense Knowledge into Story Ending Generation via Heterog…☆15May 12, 2022Updated 4 years ago
- implement byol in cifar-10☆16May 9, 2022Updated 4 years ago
- ☆17Sep 27, 2020Updated 5 years ago
- Code and Benchmarks for JOSIE (SIGMOD 2019)☆20Apr 13, 2023Updated 3 years ago
- Code for RK-Net☆32Mar 26, 2023Updated 3 years ago
- Code for "Are “Hierarchical” Visual Representations Hierarchical?" in NeurIPS Workshop for Symmetry and Geometry in Neural Representation…☆23Nov 8, 2023Updated 2 years ago
- Implemention of "Realtime Multi Person Pose-Estimation" in pytorch with data from AI Challenger☆13Nov 24, 2017Updated 8 years ago
- ☆22Sep 20, 2022Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆24Oct 9, 2023Updated 2 years ago
- We propose MMAD, a novel automated pipeline for precise AD generation. MMAD introduces ambient music alongside visual and linguistic, enh…☆17Dec 31, 2024Updated last year
- Source code of our MM'22 paper Cross-Lingual Cross-Modal Retrieval with Noise-Robust Learning☆21Jun 20, 2024Updated 2 years ago
- [ICLR 2021] Heteroskedastic and Imbalanced Deep Learning with Adaptive Regularization☆42May 2, 2021Updated 5 years ago
- Source code of WSiP model☆13Aug 14, 2022Updated 4 years ago
- GPU implementation of improved dense trajectory☆10Apr 14, 2015Updated 11 years ago
- Code for Static and Dynamic Concepts for Self-supervised Video Representation Learning.☆11Jul 28, 2022Updated 4 years ago