A curated list of awesome Multimodal studies.
β347Aug 4, 2026Updated last month
Alternatives and similar repositories for Awesome-Multimodal-Papers
Users that are interested in Awesome-Multimodal-Papers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π This is a repository for organizing papers, codes and other resources related to unified multimodal models.β833Oct 10, 2025Updated 11 months ago
- Latest Advances on Multimodal Large Language Modelsβ18,040Sep 18, 2026Updated last week
- official code for "Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval"β42Jul 4, 2025Updated last year
- Official Repository of RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuningβ15Jul 9, 2025Updated last year
- π Awesome papers on token redundancy reductionβ14Mar 12, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.β763May 21, 2026Updated 4 months ago
- An open-source implementaion for fine-tuning DINOv2 by Meta.β18Jul 21, 2025Updated last year
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-basβ¦β1,444Aug 2, 2026Updated last month
- A Survey on Benchmarks of Multimodal Large Language Modelsβ159Jul 13, 2026Updated 2 months ago
- Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a releasβ¦β1,323Jul 31, 2026Updated last month
- π₯π₯π₯ A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).β552Apr 4, 2025Updated last year
- This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]β686Sep 21, 2026Updated last week
- The codebase for our EMNLP24 paper: Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Moβ¦β85Jan 27, 2025Updated last year
- Reading list for research topics in multimodal machine learningβ6,935Aug 20, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Collection of AWESOME vision-language models for vision tasksβ3,126Sep 16, 2026Updated last week
- β21Jul 9, 2025Updated last year
- Code for "CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning"β34Mar 26, 2025Updated last year
- β¨β¨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Modelsβ43Apr 10, 2025Updated last year
- [ICLR2025] MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Modelsβ99Sep 14, 2024Updated 2 years ago
- β364Jan 27, 2024Updated 2 years ago
- [CVPR 2025 Oral] VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selectionβ142Jul 28, 2025Updated last year
- [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedbackβ311Sep 11, 2024Updated 2 years ago
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Surveyβ1,027May 22, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official repository for LLaVA-Reward (ICCV 2025): Multimodal LLMs as Customized Reward Models for Text-to-Image Generationβ27Jul 30, 2025Updated last year
- Official repo for "AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability"β34Jul 12, 2024Updated 2 years ago
- Awesome Unified Multimodal Modelsβ1,322Mar 24, 2026Updated 6 months ago
- This is the code for Multi-Behavioral Sequential Recommendation paper Accepted at RecSys 2024β11Jan 8, 2025Updated last year
- β19Oct 28, 2025Updated 11 months ago
- Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Modβ¦β378Jul 3, 2026Updated 2 months ago
- This repository is related to 'Intriguing Properties of Hyperbolic Embeddings in Vision-Language Models', published at TMLR (2024), httpsβ¦β21Jul 5, 2024Updated 2 years ago
- [ACL 2024]β61Jun 20, 2024Updated 2 years ago
- π A curated list of resources dedicated to hallucination of multimodal large language models (MLLM).β1,044Sep 27, 2025Updated last year
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [NeurIPS2024] Repo for the paper `ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models'β214Jul 17, 2025Updated last year
- A curated list of balanced multimodal learning methods.β174Mar 26, 2026Updated 6 months ago
- [TMLR 2025π₯] A survey for the autoregressive models in vision.β808May 5, 2026Updated 4 months ago
- [MM'2024] Official release of RFUND introduced in the MM'2024 paper "PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking fβ¦β21Dec 4, 2024Updated last year
- Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Modelsβ1,201Jul 14, 2026Updated 2 months ago
- [Survey] Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Surveyβ478Jan 17, 2025Updated last year
- [CVPR'24] Official implementation of our paper "Self-Supervised Facial Representation Learning with Facial Region Awareness"β15Mar 8, 2024Updated 2 years ago