[NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding
β532Nov 14, 2025Updated 10 months ago
Alternatives and similar repositories for UniTok
Users that are interested in UniTok are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR 2025] π₯ Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".β464Aug 8, 2025Updated last year
- [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representationsβ202Sep 18, 2025Updated last year
- β323May 29, 2025Updated last year
- [ICCV 2025] Official repo for "GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation"β205Jan 7, 2026Updated 8 months ago
- Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"β431Jun 20, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"β2,016Feb 25, 2026Updated 6 months ago
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,967Aug 15, 2024Updated 2 years ago
- [ICCV2025]Code Release of Harmonizing Visual Representations for Unified Multimodal Understanding and Generationβ193May 21, 2025Updated last year
- Official implementation of BLIP3o-Seriesβ1,667Nov 29, 2025Updated 9 months ago
- Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"β200Feb 24, 2026Updated 6 months ago
- This repo contains the code for 1D tokenizer and generatorβ1,175Mar 20, 2025Updated last year
- SEED-Voken: A Series of Powerful Visual Tokenizersβ1,022Nov 25, 2025Updated 9 months ago
- [ICLR 2025] VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generationβ428Apr 25, 2025Updated last year
- High-performance Image Tokenizers for VAR and ARβ307Apr 25, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Native Multimodal Models are World Learnersβ1,555Dec 30, 2025Updated 8 months ago
- [ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Thinkβ1,712Mar 16, 2025Updated last year
- Open-source unified multimodal modelβ6,176May 4, 2026Updated 4 months ago
- [CVPR 2025 Oral] Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Modelsβ1,540Dec 16, 2025Updated 9 months ago
- [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformersβ517Dec 6, 2025Updated 9 months ago
- [ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generationβ508Apr 15, 2026Updated 5 months ago
- Official Implementation of "UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation"β143Oct 17, 2025Updated 11 months ago
- (Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generatorsβ640Jun 1, 2026Updated 3 months ago
- This repository includes the official implementation of our paper "Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generatβ¦β251Oct 12, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code release for Ming-UniVision: Joint Image Understanding and Geneation with a Continuous Unified Tokenizerβ143Oct 14, 2025Updated 11 months ago
- This is a repo to track the latest autoregressive visual generation papers.β433Jun 25, 2025Updated last year
- [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.β1,977Jan 8, 2026Updated 8 months ago
- [arXiv: 2502.05178] QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generationβ97Mar 1, 2025Updated last year
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoningβ239May 30, 2025Updated last year
- FlexTok: Resampling Images into 1D Token Sequences of Flexible Lengthβ333Sep 11, 2026Updated last week
- [ICLR2026] WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstructionβ71Sep 3, 2025Updated last year
- [CVPR 2025 Oral]Infinity β : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesisβ1,588Apr 16, 2026Updated 5 months ago
- UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generationβ891Dec 23, 2025Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICML'25] EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling.β184Mar 18, 2026Updated 6 months ago
- [ICCV2025] TokenBridge: Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation. https://yuqingwang1029.github.io/Toβ¦β159Jul 24, 2025Updated last year
- PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838β1,952Feb 20, 2026Updated 6 months ago
- Official Implementation of Paper Transfer between Modalities with MetaQueriesβ327Oct 12, 2025Updated 11 months ago
- Awesome Unified Multimodal Modelsβ1,322Mar 24, 2026Updated 5 months ago
- Code for MetaMorph Multimodal Understanding and Generation via Instruction Tuningβ236Jan 22, 2026Updated 7 months ago
- MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)β1,670Feb 14, 2026Updated 7 months ago