Official implementation of SEED-LLaMA (ICLR 2024).
β642Sep 21, 2024Updated last year
Alternatives and similar repositories for SEED
Users that are interested in SEED are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Multimodal Models in Real Worldβ558Feb 24, 2025Updated last year
- π Code and models for the NeurIPS 2023 paper "Generating Images with Multimodal Language Models".β470Jan 19, 2024Updated 2 years ago
- Emu Series: Generative Multimodal Models from BAAIβ1,776Jan 12, 2026Updated 6 months ago
- LaVIT: Empower the Large Language Model to Understand and Generate Visual Contentβ603Oct 6, 2024Updated last year
- (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.β365Jan 14, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2024 Spotlight] DreamLLM: Synergistic Multimodal Comprehension and Creationβ462Dec 2, 2024Updated last year
- SEED-Voken: A Series of Powerful Visual Tokenizersβ1,017Nov 25, 2025Updated 7 months ago
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,960Aug 15, 2024Updated last year
- [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.β1,963Jan 8, 2026Updated 6 months ago
- Official code for "What Makes for Good Visual Tokenizers for Large Language Models?".β59Jun 27, 2023Updated 3 years ago
- Next-Token Prediction is All You Needβ2,432Jan 12, 2026Updated 6 months ago
- β650Feb 15, 2024Updated 2 years ago
- Official implementation of paper "MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens"β866May 8, 2025Updated last year
- MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.β953Mar 19, 2025Updated last year
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- EVE Series: Encoder-Free Vision-Language Models from BAAIβ374Jul 24, 2025Updated 11 months ago
- MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizerβ255Apr 3, 2024Updated 2 years ago
- This repo contains the code for 1D tokenizer and generatorβ1,166Mar 20, 2025Updated last year
- EVA Series: Visual Representation Fantasies from BAAIβ2,684Aug 1, 2024Updated last year
- [CVPR 2024] CapsFusion: Rethinking Image-Text Data at Scaleβ215Feb 27, 2024Updated 2 years ago
- (ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interestβ556Jun 3, 2025Updated last year
- Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.β2,104Jul 29, 2024Updated last year
- Code and models for the paper "One Transformer Fits All Distributions in Multi-Modal Diffusion"β1,486May 31, 2023Updated 3 years ago
- [CVPR 2023] Official Implementation of X-Decoder for generalized decoding for pixel, image and languageβ1,346Oct 5, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- LAVIS - A One-stop Library for Language-Vision Intelligenceβ11,255Jun 2, 2026Updated last month
- Official PyTorch implementation of the paper "In-Context Learning Unlocked for Diffusion Models"β414Mar 25, 2024Updated 2 years ago
- Implementation of MagViT2 Tokenizer in Pytorchβ668Jan 12, 2025Updated last year
- Official Implementation of "Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretrainiβ¦β646Oct 16, 2025Updated 9 months ago
- Diffusion Powers Video Tokenizer for Comprehension and Generation (CVPR 2025)β87Feb 27, 2025Updated last year
- Cambrian-1 is a family of multimodal LLMs with a vision-centric design.β2,008Nov 7, 2025Updated 8 months ago
- VisionLLM Seriesβ1,152Feb 27, 2025Updated last year
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perceptionβ159Dec 6, 2024Updated last year
- β4,710Jun 15, 2026Updated last month
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- An open-source framework for training large multimodal models.β4,114Aug 31, 2024Updated last year
- 𦦠Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing impβ¦β3,425Mar 5, 2024Updated 2 years ago
- [Extended verision ICLR 2025 Blog Track] Anole: An Open, Autoregressive and Native Multimodal Models for Interleaved Image-Text Generatioβ¦β841Jun 16, 2025Updated last year
- β814Jul 8, 2024Updated 2 years ago
- [NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.β325Jul 9, 2024Updated 2 years ago
- Densely Captioned Images (DCI) dataset repository.β197Jul 1, 2024Updated 2 years ago
- Turning to Video for Transcript Sortingβ49Aug 27, 2023Updated 2 years ago