This project aims to collect and collate various datasets for multimodal large model training, including but not limited to pre-training data, instruction fine-tuning data, and In-Context learning data.
☆78May 7, 2025Updated last year
Alternatives and similar repositories for Awesome-MLLM-Datasets
Users that are interested in Awesome-MLLM-Datasets are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20May 14, 2024Updated 2 years ago
- Fast LLM Training CodeBase With dynamic strategy choosing [Deepspeed+Megatron+FlashAttention+CudaFusionKernel+Compiler];☆42Jan 4, 2024Updated 2 years ago
- Code for paper: Unified Text-to-Image Generation and Retrieval☆15Jul 19, 2026Updated last month
- Multi-Task instruction-tuned LLaMA☆14May 5, 2023Updated 3 years ago
- Benchmarking End-to-End Photographed Document Parsing and Translation☆18Dec 4, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- MaXM is a suite of test-only benchmarks for multilingual visual question answering in 7 languages: English (en), French (fr), Hindi (hi),…☆13Jan 16, 2024Updated 2 years ago
- Paper collections of multi-modal LLM for Math/STEM/Code.☆147May 17, 2026Updated 3 months ago
- [ICLR 2024] Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement.☆15Mar 12, 2024Updated 2 years ago
- This repo offers advanced tutorials for LLMs, BERT-based models, and multimodal models, covering fine-tuning, quantization, vocabulary ex…☆24May 5, 2025Updated last year
- 🔨🔨🔨Tool for making model training data set☆20Nov 1, 2024Updated last year
- [IJCAI' 22] Uncertainty-Guided Pixel Contrastive Learning for Semi-Supervised Medical Image Segmentation.☆27Aug 4, 2022Updated 4 years ago
- paper list, tutorial, and nano code snippet for Diffusion Large Language Models.☆170Jan 19, 2026Updated 7 months ago
- GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.☆87Dec 17, 2024Updated last year
- [ICLR 2025] Code&Data for the paper "Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization"☆15Jun 21, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL 2025] Towards Text-Image Interleaved Retrieval☆16Sep 3, 2025Updated 11 months ago
- A Collection of Papers on Diffusion Language Models☆182Aug 2, 2026Updated 3 weeks ago
- Code for Cross-dataset Training☆15Dec 27, 2020Updated 5 years ago
- [ICML2024] Repo for the paper `Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models'☆24Jan 1, 2025Updated last year
- Awesome paper for multi-modal llm with grounding ability☆21Oct 11, 2025Updated 10 months ago
- ☆112Sep 11, 2025Updated 11 months ago
- 这个项目是基于python3的mxnet框架实现的实时视频人脸识别,其中包括视频传输,人脸识别等部分,用户可根据需要调整使用。整个项目建立在ubuntu18.04系统下。☆15Dec 12, 2020Updated 5 years ago
- ☆10Oct 1, 2024Updated last year
- [EMNLP'2023 Findings] MoqaGPT, for zero-shot multimodal question answering with LLMs☆13Dec 28, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- a py3 lib for NLP & image-caption metrics : BLEU METEOR CIDEr ROUGE SPICE WMD☆14Sep 13, 2022Updated 3 years ago
- [ACL 2025] Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging☆40Jun 4, 2025Updated last year
- [IEEE TMM'25] Scene-Text Grounding for Text-Based Video Question Answering☆17Feb 16, 2026Updated 6 months ago
- Advancing TTP Analysis: Harnessing the Power of Large Language Models with Retrieval Augmented Generation☆11May 14, 2024Updated 2 years ago
- In this playground competition, you are provided a strictly canine subset of ImageNet in order to practice fine-grained image categorizat…☆11Dec 10, 2020Updated 5 years ago
- We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing their…☆22Jan 11, 2026Updated 7 months ago
- Code for "RSF: Optimizing Rigid Scene Flow From 3D Point Clouds Without Labels"☆10Jan 17, 2023Updated 3 years ago
- 基于文字密度的新闻正文提取模块,兼容python2和python3,传入新闻网址或者网页源码即可返回标题,发布时间和正文内容。☆14Jun 10, 2018Updated 8 years ago
- [ACL2025 main] Official implementation of "LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjo…☆20Mar 16, 2026Updated 5 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Codes for "Benchmarking the Generation of Fact Checking Explanations"☆10Aug 16, 2024Updated 2 years ago
- ☆12Feb 2, 2023Updated 3 years ago
- Data-Efficient Multimodal Fusion on a Single GPU☆68May 7, 2024Updated 2 years ago
- Source code for AAAI'25 paper "Component-Level Segmentation for Oracle Bone Inscription Decipherment"☆20Oct 13, 2025Updated 10 months ago
- ☆20Mar 5, 2025Updated last year
- Code release for "Weakly Supervised Open-Vocabulary Object Detection", AAAI2024☆36Sep 9, 2024Updated last year
- ☆52May 28, 2024Updated 2 years ago