This project aims to collect and collate various datasets for multimodal large model training, including but not limited to pre-training data, instruction fine-tuning data, and In-Context learning data.
☆78May 7, 2025Updated last year
Alternatives and similar repositories for Awesome-MLLM-Datasets
Users that are interested in Awesome-MLLM-Datasets are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20May 14, 2024Updated 2 years ago
- Fast LLM Training CodeBase With dynamic strategy choosing [Deepspeed+Megatron+FlashAttention+CudaFusionKernel+Compiler];☆42Jan 4, 2024Updated 2 years ago
- Federated Meta-Learning for Emotion and Sentiment Aware Multi-modal Complaint Identification☆10May 30, 2024Updated 2 years ago
- Code for paper: Unified Text-to-Image Generation and Retrieval☆15Jul 19, 2026Updated 2 months ago
- Multi-Task instruction-tuned LLaMA☆14May 5, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [CVPR 2025] Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding☆17Jun 16, 2025Updated last year
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-bas…☆1,441Aug 2, 2026Updated last month
- Paper collections of multi-modal LLM for Math/STEM/Code.☆146Aug 30, 2026Updated 3 weeks ago
- [ICLR 2024] Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement.☆15Mar 12, 2024Updated 2 years ago
- Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Mod…☆378Jul 3, 2026Updated 2 months ago
- 从socket开始实现pop3和smtp客户端,实现邮件编写、发送、接收、阅读、删除等基本功能。并实现简单界面(PyQt5)Start from socket to implement pop3 and smtp clients, to realize the basic …☆12Dec 24, 2023Updated 2 years ago
- Efficient Segment Anything in Medical Images☆46Jul 27, 2024Updated 2 years ago
- 🔨🔨🔨Tool for making model training data set☆20Nov 1, 2024Updated last year
- ☆26Feb 2, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [IJCAI' 22] Uncertainty-Guided Pixel Contrastive Learning for Semi-Supervised Medical Image Segmentation.☆27Aug 4, 2022Updated 4 years ago
- Using convolutional neural networks for the 2019 Kidney and Kidney Tumor Segmentation Challenge☆19Dec 13, 2019Updated 6 years ago
- Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective (ACL 2024)☆59Oct 28, 2024Updated last year
- PyTorch implementation of "UNIT: Unifying Image and Text Recognition in One Vision Encoder", NeurlPS 2024.☆34Sep 26, 2024Updated last year
- [ICLR 2025] Code&Data for the paper "Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization"☆15Jun 21, 2024Updated 2 years ago
- [EMNLP 2022] Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive Learning☆14Jun 22, 2024Updated 2 years ago
- iSegFormer: Interactive Image/Volume Segmentation using Vision Transformers (MICCAI 2022)☆31Oct 24, 2025Updated 10 months ago
- A Collection of Papers on Diffusion Language Models☆186Aug 2, 2026Updated last month
- [ICML2024] Repo for the paper `Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models'☆24Jan 1, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆111Sep 11, 2025Updated last year
- ☆10Oct 1, 2024Updated last year
- Top 3 solution for CVPR24 SEGMENT ANYTHING IN MEDICAL IMAGES ON LAPTOP Challenge☆11Apr 8, 2025Updated last year
- This repository will continuously update the latest papers, technical reports, benchmarks about multimodal reasoning!☆56Mar 21, 2025Updated last year
- [EMNLP'2023 Findings] MoqaGPT, for zero-shot multimodal question answering with LLMs☆13Dec 28, 2024Updated last year
- Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts☆15Feb 26, 2024Updated 2 years ago
- a py3 lib for NLP & image-caption metrics : BLEU METEOR CIDEr ROUGE SPICE WMD☆14Sep 13, 2022Updated 4 years ago
- [IEEE TMM'25] Scene-Text Grounding for Text-Based Video Question Answering☆17Feb 16, 2026Updated 7 months ago
- In this playground competition, you are provided a strictly canine subset of ImageNet in order to practice fine-grained image categorizat…☆11Dec 10, 2020Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing their…☆23Jan 11, 2026Updated 8 months ago
- Codes for "Benchmarking the Generation of Fact Checking Explanations"☆11Aug 16, 2024Updated 2 years ago
- ☆12Feb 2, 2023Updated 3 years ago
- [ACL 2025] Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence☆20Jun 10, 2025Updated last year
- ☆20Mar 5, 2025Updated last year
- Code release for "Weakly Supervised Open-Vocabulary Object Detection", AAAI2024☆36Sep 9, 2024Updated 2 years ago
- ☆15Jul 14, 2022Updated 4 years ago