A curated list of awesome Multimodal studies.
β344Aug 4, 2026Updated 2 weeks ago
Alternatives and similar repositories for Awesome-Multimodal-Papers
Users that are interested in Awesome-Multimodal-Papers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π This is a repository for organizing papers, codes and other resources related to unified multimodal models.β827Oct 10, 2025Updated 10 months ago
- Latest Advances on Multimodal Large Language Modelsβ17,981Updated this week
- official code for "Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval"β42Jul 4, 2025Updated last year
- Official Repository of RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuningβ14Jul 9, 2025Updated last year
- π Awesome papers on token redundancy reductionβ14Mar 12, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.β761May 21, 2026Updated 2 months ago
- An open-source implementaion for fine-tuning DINOv2 by Meta.β16Jul 21, 2025Updated last year
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-basβ¦β1,441Aug 2, 2026Updated 2 weeks ago
- A Survey on Benchmarks of Multimodal Large Language Modelsβ158Jul 13, 2026Updated last month
- Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a releasβ¦β1,304Jul 31, 2026Updated 2 weeks ago
- π₯π₯π₯ A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).β552Apr 4, 2025Updated last year
- This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]β675Jul 24, 2026Updated 3 weeks ago
- The codebase for our EMNLP24 paper: Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Moβ¦β85Jan 27, 2025Updated last year
- Reading list for research topics in multimodal machine learningβ6,925Aug 20, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Collection of AWESOME vision-language models for vision tasksβ3,128Oct 14, 2025Updated 10 months ago
- β21Jul 9, 2025Updated last year
- Code for "CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning"β33Mar 26, 2025Updated last year
- β¨β¨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Modelsβ43Apr 10, 2025Updated last year
- [ICLR2025] MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Modelsβ98Sep 14, 2024Updated last year
- β365Jan 27, 2024Updated 2 years ago
- [CVPR 2025 Oral] VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selectionβ141Jul 28, 2025Updated last year
- [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedbackβ311Sep 11, 2024Updated last year
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Surveyβ1,018May 22, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official repository for LLaVA-Reward (ICCV 2025): Multimodal LLMs as Customized Reward Models for Text-to-Image Generationβ26Jul 30, 2025Updated last year
- Official repo for "AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability"β34Jul 12, 2024Updated 2 years ago
- Awesome Unified Multimodal Modelsβ1,311Mar 24, 2026Updated 4 months ago
- This is the code for Multi-Behavioral Sequential Recommendation paper Accepted at RecSys 2024β12Jan 8, 2025Updated last year
- β19Oct 28, 2025Updated 9 months ago
- Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Modβ¦β377Jul 3, 2026Updated last month
- This repository is related to 'Intriguing Properties of Hyperbolic Embeddings in Vision-Language Models', published at TMLR (2024), httpsβ¦β21Jul 5, 2024Updated 2 years ago
- [ACL 2024]β60Jun 20, 2024Updated 2 years ago
- π A curated list of resources dedicated to hallucination of multimodal large language models (MLLM).β1,038Sep 27, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [NeurIPS2024] Repo for the paper `ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models'β211Jul 17, 2025Updated last year
- [TMLR 2025π₯] A survey for the autoregressive models in vision.β806May 5, 2026Updated 3 months ago
- A curated list of balanced multimodal learning methods.β172Mar 26, 2026Updated 4 months ago
- Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Modelsβ1,193Jul 14, 2026Updated last month
- [Survey] Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Surveyβ477Jan 17, 2025Updated last year
- [CVPR'24] Official implementation of our paper "Self-Supervised Facial Representation Learning with Facial Region Awareness"β15Mar 8, 2024Updated 2 years ago
- β46Dec 30, 2024Updated last year