Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"
β431Jun 20, 2025Updated last year
Alternatives and similar repositories for SimpleAR
Users that are interested in SimpleAR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICCV 2025] Official repo for "GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation"β204Jan 7, 2026Updated 6 months ago
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,959Aug 15, 2024Updated last year
- [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understandingβ529Nov 14, 2025Updated 8 months ago
- [ICLR 2025] Autoregressive Video Generation without Vector Quantizationβ656Oct 29, 2025Updated 8 months ago
- [NeurIPS 2025] T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTβ433Sep 18, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2025 Oral]Infinity β : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesisβ1,579Apr 16, 2026Updated 3 months ago
- SEED-Voken: A Series of Powerful Visual Tokenizersβ1,016Nov 25, 2025Updated 7 months ago
- [ICCV 2025] The official implementation of "Neighboring Autoregressive Modeling for Efficient Visual Generation"β62Apr 5, 2025Updated last year
- [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.β1,963Jan 8, 2026Updated 6 months ago
- β21Jan 17, 2025Updated last year
- [ICCV2025]Code Release of Harmonizing Visual Representations for Unified Multimodal Understanding and Generationβ191May 21, 2025Updated last year
- [CVPR 2025] π₯ Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".β464Aug 8, 2025Updated 11 months ago
- Official Implementation of Paper Transfer between Modalities with MetaQueriesβ324Oct 12, 2025Updated 9 months ago
- PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838β1,942Feb 20, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- HART: Efficient Visual Generation with Hybrid Autoregressive Transformerβ647Oct 16, 2024Updated last year
- Official implementation of BLIP3o-Seriesβ1,663Nov 29, 2025Updated 7 months ago
- [CVPR 2025] The First Investigation of CoT Reasoning (RL, TTS, Reflection) in Image Generationβ865Mar 19, 2026Updated 4 months ago
- [NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RLβ2,420May 7, 2026Updated 2 months ago
- This repo contains the code for 1D tokenizer and generatorβ1,165Mar 20, 2025Updated last year
- Open-source unified multimodal modelβ6,103May 4, 2026Updated 2 months ago
- Official Implementation of "Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretrainiβ¦β646Oct 16, 2025Updated 9 months ago
- UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generationβ883Dec 23, 2025Updated 6 months ago
- VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learningβ271Apr 15, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representationsβ202Sep 18, 2025Updated 10 months ago
- [ICLR 2026] This is an early exploration to introduce Interleaving Reasoning to Text-to-image Generation field and achieve the SoTA benchβ¦β100Jan 26, 2026Updated 5 months ago
- Pixel-Space Generative Modelsβ316May 11, 2025Updated last year
- Official repository of "GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing"β317Sep 28, 2025Updated 9 months ago
- Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).β426Aug 26, 2025Updated 10 months ago
- [CVPR2025] Official Implementation of ILLUME+β126Aug 20, 2025Updated 11 months ago
- [ICCV2025] TokenBridge: Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation. https://yuqingwang1029.github.io/Toβ¦β158Jul 24, 2025Updated 11 months ago
- [CVPR2025 Highlight] PAR: Parallelized Autoregressive Visual Generation. https://yuqingwang1029.github.io/PAR-projectβ186Mar 20, 2025Updated last year
- [CVPR 2025 Oral] Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Modelsβ1,507Dec 16, 2025Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ARM: An AutoRegressive Large Multimodal Model with Discrete Representationsβ50Jun 10, 2026Updated last month
- Native Multimodal Models are World Learnersβ1,536Dec 30, 2025Updated 6 months ago
- β25Dec 26, 2024Updated last year
- [ICLR 2025] VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generationβ425Apr 25, 2025Updated last year
- [ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potentiβ¦β410May 23, 2026Updated last month
- Official implementation of HPSv3: Towards Wide-Spectrum Human Preference Score (ICCV2025)β325Dec 5, 2025Updated 7 months ago
- MAGI-1: Autoregressive Video Generation at Scaleβ3,741Jun 17, 2026Updated last month