ARM: An AutoRegressive Large Multimodal Model with Discrete Representations
☆50Jun 10, 2026Updated last month
Alternatives and similar repositories for ARM
Users that are interested in ARM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key …☆49Jul 13, 2026Updated 3 weeks ago
- ☆25Dec 26, 2024Updated last year
- ☆21Jan 17, 2025Updated last year
- Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers☆60Jun 14, 2026Updated last month
- [ICML 2026] VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding☆27Jul 3, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS 2025] The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tun…☆40Feb 20, 2025Updated last year
- [CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"☆19Mar 9, 2026Updated 4 months ago
- [CVPR 2026] FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding☆75Mar 16, 2026Updated 4 months ago
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation☆43Nov 24, 2025Updated 8 months ago
- Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"☆431Jun 20, 2025Updated last year
- A controllable and interactive simulation framework for vision research.☆16May 25, 2026Updated 2 months ago
- [CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"☆52Jun 16, 2025Updated last year
- A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing☆20Mar 13, 2026Updated 4 months ago
- [CVPR 2026] FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance☆66Mar 13, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- FNIN: A Fourier Neural Operator-based Numerical Integration Network for Surface-form-gradients☆13Jan 22, 2025Updated last year
- [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding☆530Nov 14, 2025Updated 8 months ago
- ☆58Jun 4, 2024Updated 2 years ago
- This repository provides the official implementation of VTBench, a benchmark designed to evaluate the performance of visual tokenizers (V…☆35Jul 30, 2025Updated last year
- [AAAI-25] Official repository of "Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object De…☆20Dec 27, 2024Updated last year
- Official implementation for NeurIPS'25 paper "NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses"☆21Nov 21, 2025Updated 8 months ago
- [ICCV 2025] MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance☆184Feb 11, 2026Updated 5 months ago
- Official PyTorch implementation for "Effective and Efficient Masked Image Generation Models"☆35Apr 8, 2025Updated last year
- [SIGGRAPH Asia 2026] DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models☆152Jul 20, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2026] Lumos Project: Frontier video unified model research by Alibaba DAMO Academy.☆161Apr 6, 2026Updated 3 months ago
- [CVPR2026 Highlight] Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens https://arxiv.org/abs…☆63Apr 10, 2026Updated 3 months ago
- ☆20Dec 8, 2024Updated last year
- ☆28Apr 4, 2025Updated last year
- Decoupled Memory Selection for Multi-target Video Segmentation of SAM3☆59Jan 16, 2026Updated 6 months ago
- [NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.☆325Jul 9, 2024Updated 2 years ago
- [ICLR 2026] Official code for "Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks"☆27Mar 2, 2026Updated 5 months ago
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation☆38Sep 16, 2025Updated 10 months ago
- Next Forcing: World Action Modeling with Multi-Chunk Prediction (MCP)☆114Jul 19, 2026Updated 2 weeks ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICLR 2026] This is an early exploration to introduce Interleaving Reasoning to Text-to-image Generation field and achieve the SoTA bench…☆100Jan 26, 2026Updated 6 months ago
- PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokens☆18Nov 26, 2025Updated 8 months ago
- Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation☆739Jul 22, 2026Updated last week
- [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations☆202Sep 18, 2025Updated 10 months ago
- ☆39Dec 4, 2023Updated 2 years ago
- ☆20Apr 2, 2026Updated 4 months ago
- ☆69May 14, 2026Updated 2 months ago