[ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
β430May 5, 2025Updated last year
Alternatives and similar repositories for OmniCorpus
Users that are interested in OmniCorpus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π MINT-1T: A one trillion token multimodal interleaved dataset.β835Jul 31, 2024Updated 2 years ago
- [NeurIPS 2024] Needle In A Multimodal Haystack (MM-NIAH): A comprehensive benchmark designed to systematically evaluate the capability ofβ¦β130Nov 25, 2024Updated last year
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perceptionβ161Dec 6, 2024Updated last year
- [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of β¦β507Aug 9, 2024Updated 2 years ago
- EVE Series: Encoder-Free Vision-Language Models from BAAIβ376Jul 24, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Modelβ282Jun 25, 2024Updated 2 years ago
- [CVPR 2024] CapsFusion: Rethinking Image-Text Data at Scaleβ215Feb 27, 2024Updated 2 years ago
- [NeurIPS 2024] Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learningβ72Feb 11, 2025Updated last year
- Cambrian-1 is a family of multimodal LLMs with a vision-centric design.β2,011Nov 7, 2025Updated 11 months ago
- Learning 1D Causal Visual Representation with De-focus Attention Networksβ35Jun 7, 2024Updated 2 years ago
- [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. ζ₯θΏGPT-4o葨η°ηεΌζΊε€ζ¨‘ζε―Ήθ―樑εβ10,166Sep 22, 2025Updated last year
- β134Dec 22, 2023Updated 2 years ago
- LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMsβ426Jul 6, 2026Updated 3 months ago
- MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizerβ255Apr 3, 2024Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuningβ297Mar 13, 2024Updated 2 years ago
- [CVPR'25 highlight] RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthinessβ461May 14, 2025Updated last year
- LaVIT: Empower the Large Language Model to Understand and Generate Visual Contentβ606Oct 6, 2024Updated 2 years ago
- Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarksβ4,436Updated this week
- Emu Series: Generative Multimodal Models from BAAIβ1,777Jan 12, 2026Updated 8 months ago
- [ICCV 2025 Highlight] The official repository for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining"β197Mar 17, 2025Updated last year
- [CVPR 2025 Highlight] Official repository for CoMM Datasetβ59Dec 31, 2024Updated last year
- [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedbackβ310Sep 11, 2024Updated 2 years ago
- Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.β2,102Jul 29, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- VisionLLM Seriesβ1,154Feb 27, 2025Updated last year
- Official repository of MMDU datasetβ110Sep 29, 2024Updated 2 years ago
- Code used for the creation of OBELICS, an open, massive and curated collection of interleaved image-text web documents, containing 141M dβ¦β219Aug 28, 2024Updated 2 years ago
- When do we not need larger vision models?β418Feb 8, 2025Updated last year
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving statβ¦β1,593Jun 14, 2025Updated last year
- EVA Series: Visual Representation Fantasies from BAAIβ2,694Aug 1, 2024Updated 2 years ago
- β¨β¨Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Modelsβ161Dec 26, 2024Updated last year
- A huge dataset for Document Visual Question Answeringβ24Jul 29, 2024Updated 2 years ago
- A fork to add multimodal model training to open-r1β1,605Feb 8, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactionsβ2,928May 26, 2025Updated last year
- β4,729Jun 15, 2026Updated 3 months ago
- NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024β1,854Aug 11, 2026Updated last month
- SVIT: Scaling up Visual Instruction Tuningβ168Jun 20, 2024Updated 2 years ago
- Official repo for paper "MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions"β530Sep 2, 2024Updated 2 years ago
- [ICML 2025] This is the official repository of our paper "What If We Recaption Billions of Web Images with LLaMA-3 ?"β152Jun 13, 2024Updated 2 years ago
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,968Aug 15, 2024Updated 2 years ago