[ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
β426May 5, 2025Updated last year
Alternatives and similar repositories for OmniCorpus
Users that are interested in OmniCorpus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π MINT-1T: A one trillion token multimodal interleaved dataset.β832Jul 31, 2024Updated 2 years ago
- [NeurIPS 2024] Needle In A Multimodal Haystack (MM-NIAH): A comprehensive benchmark designed to systematically evaluate the capability ofβ¦β127Nov 25, 2024Updated last year
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perceptionβ159Dec 6, 2024Updated last year
- [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of β¦β506Aug 9, 2024Updated 2 years ago
- EVE Series: Encoder-Free Vision-Language Models from BAAIβ376Jul 24, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Modelβ281Jun 25, 2024Updated 2 years ago
- [CVPR 2024] CapsFusion: Rethinking Image-Text Data at Scaleβ215Feb 27, 2024Updated 2 years ago
- [NeurIPS 2024] Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learningβ72Feb 11, 2025Updated last year
- Cambrian-1 is a family of multimodal LLMs with a vision-centric design.β2,014Nov 7, 2025Updated 9 months ago
- Learning 1D Causal Visual Representation with De-focus Attention Networksβ35Jun 7, 2024Updated 2 years ago
- [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. ζ₯θΏGPT-4o葨η°ηεΌζΊε€ζ¨‘ζε―Ήθ―樑εβ10,120Sep 22, 2025Updated 10 months ago
- β134Dec 22, 2023Updated 2 years ago
- LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMsβ423Jul 6, 2026Updated last month
- MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizerβ256Apr 3, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [CVPR'25 highlight] RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthinessβ459May 14, 2025Updated last year
- [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuningβ297Mar 13, 2024Updated 2 years ago
- LaVIT: Empower the Large Language Model to Understand and Generate Visual Contentβ604Oct 6, 2024Updated last year
- Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarksβ4,336Updated this week
- Emu Series: Generative Multimodal Models from BAAIβ1,777Jan 12, 2026Updated 6 months ago
- [ICCV 2025 Highlight] The official repository for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining"β198Mar 17, 2025Updated last year
- [CVPR 2025 Highlight] Official repository for CoMM Datasetβ57Dec 31, 2024Updated last year
- [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedbackβ310Sep 11, 2024Updated last year
- Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.β2,104Jul 29, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- VisionLLM Seriesβ1,153Feb 27, 2025Updated last year
- Official repository of MMDU datasetβ109Sep 29, 2024Updated last year
- Code used for the creation of OBELICS, an open, massive and curated collection of interleaved image-text web documents, containing 141M dβ¦β217Aug 28, 2024Updated last year
- When do we not need larger vision models?β419Feb 8, 2025Updated last year
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving statβ¦β1,583Jun 14, 2025Updated last year
- EVA Series: Visual Representation Fantasies from BAAIβ2,688Aug 1, 2024Updated 2 years ago
- β¨β¨Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Modelsβ163Dec 26, 2024Updated last year
- A huge dataset for Document Visual Question Answeringβ24Jul 29, 2024Updated 2 years ago
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactionsβ2,924May 26, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A fork to add multimodal model training to open-r1β1,599Feb 8, 2025Updated last year
- NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024β1,851Nov 27, 2025Updated 8 months ago
- β4,711Jun 15, 2026Updated last month
- SVIT: Scaling up Visual Instruction Tuningβ167Jun 20, 2024Updated 2 years ago
- Official repo for paper "MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions"β527Sep 2, 2024Updated last year
- [ICML 2025] This is the official repository of our paper "What If We Recaption Billions of Web Images with LLaMA-3 ?"β152Jun 13, 2024Updated 2 years ago
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,965Aug 15, 2024Updated last year