ARM: An AutoRegressive Large Multimodal Model with Discrete Representations
☆50Jun 10, 2026Updated 3 months ago
Alternatives and similar repositories for ARM
Users that are interested in ARM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key …☆57Jul 13, 2026Updated 2 months ago
- ☆25Dec 26, 2024Updated last year
- ☆21Jan 17, 2025Updated last year
- Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers☆66Aug 4, 2026Updated last month
- [NeurIPS 2025] The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tun…☆40Feb 20, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"☆19Mar 9, 2026Updated 6 months ago
- [CVPR 2026] FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding☆80Mar 16, 2026Updated 5 months ago
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation☆44Updated this week
- Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"☆431Jun 20, 2025Updated last year
- A controllable and interactive simulation framework for vision research.☆16May 25, 2026Updated 3 months ago
- [CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"☆53Jun 16, 2025Updated last year
- A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing☆22Mar 13, 2026Updated 6 months ago
- [CVPR 2026] FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance☆71Mar 13, 2026Updated 6 months ago
- FNIN: A Fourier Neural Operator-based Numerical Integration Network for Surface-form-gradients☆13Jan 22, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding☆529Nov 14, 2025Updated 9 months ago
- ☆58Jun 4, 2024Updated 2 years ago
- This repository provides the official implementation of VTBench, a benchmark designed to evaluate the performance of visual tokenizers (V…☆36Jul 30, 2025Updated last year
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation (CVPR'25)☆21Apr 2, 2025Updated last year
- [AAAI-25] Official repository of "Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object De…☆21Dec 27, 2024Updated last year
- Official implementation for NeurIPS'25 paper "NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses"☆21Nov 21, 2025Updated 9 months ago
- [ICCV 2025] MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance☆187Feb 11, 2026Updated 7 months ago
- Official PyTorch implementation for "Effective and Efficient Masked Image Generation Models"☆35Apr 8, 2025Updated last year
- [SIGGRAPH Asia 2026] DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models☆176Jul 20, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2026] Lumos Project: Frontier video unified model research by Alibaba DAMO Academy.☆261Apr 6, 2026Updated 5 months ago
- [CVPR2026 Highlight] Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens https://arxiv.org/abs…☆63Apr 10, 2026Updated 5 months ago
- ☆20Dec 8, 2024Updated last year
- ☆29Apr 4, 2025Updated last year
- Decoupled Memory Selection for Multi-target Video Segmentation of SAM3☆62Jan 16, 2026Updated 7 months ago
- 🔥Multi-View Subject-Consistent Video Generation (SIGGRAPH 2026)☆18Jun 13, 2026Updated 3 months ago
- code for downloading videos from HowTo100M dataset☆18May 13, 2021Updated 5 years ago
- [ICLR 2026] Official code for "Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks"☆28Mar 2, 2026Updated 6 months ago
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation☆38Sep 16, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2026] This is an early exploration to introduce Interleaving Reasoning to Text-to-image Generation field and achieve the SoTA bench…☆100Jan 26, 2026Updated 7 months ago
- Next Forcing: Causal World Modeling with Multi-Chunk Prediction (MCP)☆132Aug 16, 2026Updated 3 weeks ago
- PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokens☆18Nov 26, 2025Updated 9 months ago
- Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation☆755Jul 22, 2026Updated last month
- [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations☆202Sep 18, 2025Updated 11 months ago
- ☆39Dec 4, 2023Updated 2 years ago
- ☆22Apr 2, 2026Updated 5 months ago