[ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
β427May 5, 2025Updated last year
Alternatives and similar repositories for OmniCorpus
Users that are interested in OmniCorpus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π MINT-1T: A one trillion token multimodal interleaved dataset.β832Jul 31, 2024Updated 2 years ago
- [NeurIPS 2024] Needle In A Multimodal Haystack (MM-NIAH): A comprehensive benchmark designed to systematically evaluate the capability ofβ¦β129Nov 25, 2024Updated last year
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perceptionβ160Dec 6, 2024Updated last year
- [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of β¦β506Aug 9, 2024Updated 2 years ago
- EVE Series: Encoder-Free Vision-Language Models from BAAIβ376Jul 24, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Modelβ282Jun 25, 2024Updated 2 years ago
- [CVPR 2024] CapsFusion: Rethinking Image-Text Data at Scaleβ216Feb 27, 2024Updated 2 years ago
- [NeurIPS 2024] Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learningβ72Feb 11, 2025Updated last year
- Cambrian-1 is a family of multimodal LLMs with a vision-centric design.β2,014Nov 7, 2025Updated 9 months ago
- Learning 1D Causal Visual Representation with De-focus Attention Networksβ35Jun 7, 2024Updated 2 years ago
- [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. ζ₯θΏGPT-4o葨η°ηεΌζΊε€ζ¨‘ζε―Ήθ―樑εβ10,148Sep 22, 2025Updated 11 months ago
- β134Dec 22, 2023Updated 2 years ago
- LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMsβ424Jul 6, 2026Updated last month
- MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizerβ257Apr 3, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CVPR'25 highlight] RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthinessβ459May 14, 2025Updated last year
- [ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuningβ296Mar 13, 2024Updated 2 years ago
- LaVIT: Empower the Large Language Model to Understand and Generate Visual Contentβ604Oct 6, 2024Updated last year
- Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarksβ4,363Updated this week
- Emu Series: Generative Multimodal Models from BAAIβ1,779Jan 12, 2026Updated 7 months ago
- [ICCV 2025 Highlight] The official repository for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining"β198Mar 17, 2025Updated last year
- [CVPR 2025 Highlight] Official repository for CoMM Datasetβ59Dec 31, 2024Updated last year
- [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedbackβ311Sep 11, 2024Updated last year
- Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.β2,103Jul 29, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- VisionLLM Seriesβ1,154Feb 27, 2025Updated last year
- Official repository of MMDU datasetβ109Sep 29, 2024Updated last year
- Code used for the creation of OBELICS, an open, massive and curated collection of interleaved image-text web documents, containing 141M dβ¦β218Aug 28, 2024Updated 2 years ago
- When do we not need larger vision models?β419Feb 8, 2025Updated last year
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving statβ¦β1,586Jun 14, 2025Updated last year
- EVA Series: Visual Representation Fantasies from BAAIβ2,693Aug 1, 2024Updated 2 years ago
- β¨β¨Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Modelsβ161Dec 26, 2024Updated last year
- A huge dataset for Document Visual Question Answeringβ24Jul 29, 2024Updated 2 years ago
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactionsβ2,926May 26, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A fork to add multimodal model training to open-r1β1,603Feb 8, 2025Updated last year
- NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024β1,854Aug 11, 2026Updated 2 weeks ago
- β4,717Jun 15, 2026Updated 2 months ago
- SVIT: Scaling up Visual Instruction Tuningβ168Jun 20, 2024Updated 2 years ago
- Official repo for paper "MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions"β529Sep 2, 2024Updated last year
- [ICML 2025] This is the official repository of our paper "What If We Recaption Billions of Web Images with LLaMA-3 ?"β152Jun 13, 2024Updated 2 years ago
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,966Aug 15, 2024Updated 2 years ago