[CVPR 2025] π₯ Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".
β464Aug 8, 2025Updated 11 months ago
Alternatives and similar repositories for TokenFlow
Users that are interested in TokenFlow are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understandingβ529Nov 14, 2025Updated 8 months ago
- [ICLR 2025] VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generationβ425Apr 25, 2025Updated last year
- SEED-Voken: A Series of Powerful Visual Tokenizersβ1,016Nov 25, 2025Updated 7 months ago
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,959Aug 15, 2024Updated last year
- High-performance Image Tokenizers for VAR and ARβ307Apr 25, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- π This is a repository for organizing papers, codes and other resources related to unified multimodal models.β828Oct 10, 2025Updated 9 months ago
- [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.β1,962Jan 8, 2026Updated 6 months ago
- [ICCV 2025] Official repo for "GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation"β204Jan 7, 2026Updated 6 months ago
- [TMLR 2025π₯] A survey for the autoregressive models in vision.β804May 5, 2026Updated 2 months ago
- This repo contains the code for 1D tokenizer and generatorβ1,165Mar 20, 2025Updated last year
- [ICLR 2025] Autoregressive Video Generation without Vector Quantizationβ655Oct 29, 2025Updated 8 months ago
- Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"β431Jun 20, 2025Updated last year
- [CVPR 2025 Oral]Infinity β : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesisβ1,579Apr 16, 2026Updated 3 months ago
- PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838β1,942Feb 20, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- NextFlowπ: Unified Sequential Modeling Activates Multimodal Understanding and Generationβ331Jan 9, 2026Updated 6 months ago
- Official Implementation of "Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretrainiβ¦β646Oct 16, 2025Updated 9 months ago
- Official implementation of BLIP3o-Seriesβ1,663Nov 29, 2025Updated 7 months ago
- [ICLR 2025] ControlAR: Controllable Image Generation with Autoregressive Modelsβ326Jun 30, 2026Updated 2 weeks ago
- [CVPR2025 Highlight] PAR: Parallelized Autoregressive Visual Generation. https://yuqingwang1029.github.io/PAR-projectβ186Mar 20, 2025Updated last year
- Next-Token Prediction is All You Needβ2,432Jan 12, 2026Updated 6 months ago
- Native Multimodal Models are World Learnersβ1,535Dec 30, 2025Updated 6 months ago
- (Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generatorsβ642Jun 1, 2026Updated last month
- HART: Efficient Visual Generation with Hybrid Autoregressive Transformerβ647Oct 16, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- EVE Series: Encoder-Free Vision-Language Models from BAAIβ374Jul 24, 2025Updated 11 months ago
- This is a repo to track the latest autoregressive visual generation papers.β431Jun 25, 2025Updated last year
- FlexTok: Resampling Images into 1D Token Sequences of Flexible Lengthβ321Jun 2, 2025Updated last year
- This is the official implementation for ControlVAR.β128Dec 10, 2024Updated last year
- β196Dec 17, 2024Updated last year
- Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"β1,977Feb 25, 2026Updated 4 months ago
- This repository includes the official implementation of our paper "Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generatβ¦β251Oct 12, 2025Updated 9 months ago
- Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"β195Feb 24, 2026Updated 4 months ago
- Open-source unified multimodal modelβ6,100May 4, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [Survey] Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Surveyβ476Jan 17, 2025Updated last year
- Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAIβ1,382Jan 27, 2026Updated 5 months ago
- Implements VAR+CLIP for text-to-image (T2I) generationβ147Jan 23, 2025Updated last year
- Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).β426Aug 26, 2025Updated 10 months ago
- [CVPR 2025 (Oral)] Open implementation of "RandAR"β208Jul 14, 2025Updated last year
- Multimodal Models in Real Worldβ558Feb 24, 2025Updated last year
- [arXiv: 2502.05178] QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generationβ97Mar 1, 2025Updated last year