ARM: An AutoRegressive Large Multimodal Model with Discrete Representations
☆50Jun 10, 2026Updated 2 months ago
Alternatives and similar repositories for ARM
Users that are interested in ARM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key …☆53Jul 13, 2026Updated last month
- ☆25Dec 26, 2024Updated last year
- ☆21Jan 17, 2025Updated last year
- Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers☆64Aug 4, 2026Updated 2 weeks ago
- [ICML 2026] VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding☆27Aug 10, 2026Updated last week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [NeurIPS 2025] The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tun…☆40Feb 20, 2025Updated last year
- [CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"☆19Mar 9, 2026Updated 5 months ago
- [CVPR 2026] FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding☆76Mar 16, 2026Updated 5 months ago
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation☆43Nov 24, 2025Updated 9 months ago
- Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"☆431Jun 20, 2025Updated last year
- A controllable and interactive simulation framework for vision research.☆16May 25, 2026Updated 2 months ago
- [CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"☆52Jun 16, 2025Updated last year
- A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing☆20Mar 13, 2026Updated 5 months ago
- [CVPR 2026] FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance☆70Mar 13, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- FNIN: A Fourier Neural Operator-based Numerical Integration Network for Surface-form-gradients☆13Jan 22, 2025Updated last year
- [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding☆530Nov 14, 2025Updated 9 months ago
- ☆58Jun 4, 2024Updated 2 years ago
- This repository provides the official implementation of VTBench, a benchmark designed to evaluate the performance of visual tokenizers (V…☆36Jul 30, 2025Updated last year
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation (CVPR'25)☆21Apr 2, 2025Updated last year
- [AAAI-25] Official repository of "Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object De…☆21Dec 27, 2024Updated last year
- Official implementation for NeurIPS'25 paper "NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses"☆21Nov 21, 2025Updated 9 months ago
- [ICCV 2025] MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance☆187Feb 11, 2026Updated 6 months ago
- Official PyTorch implementation for "Effective and Efficient Masked Image Generation Models"☆35Apr 8, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [SIGGRAPH Asia 2026] DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models☆160Jul 20, 2026Updated last month
- [ICLR 2026] Lumos Project: Frontier video unified model research by Alibaba DAMO Academy.☆227Apr 6, 2026Updated 4 months ago
- [CVPR2026 Highlight] Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens https://arxiv.org/abs…☆63Apr 10, 2026Updated 4 months ago
- ☆20Dec 8, 2024Updated last year
- ☆28Apr 4, 2025Updated last year
- Decoupled Memory Selection for Multi-target Video Segmentation of SAM3☆62Jan 16, 2026Updated 7 months ago
- [NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.☆326Jul 9, 2024Updated 2 years ago
- code for downloading videos from HowTo100M dataset☆18May 13, 2021Updated 5 years ago
- [ICLR 2026] Official code for "Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks"☆28Mar 2, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation☆38Sep 16, 2025Updated 11 months ago
- [ICLR 2026] This is an early exploration to introduce Interleaving Reasoning to Text-to-image Generation field and achieve the SoTA bench…☆101Jan 26, 2026Updated 6 months ago
- Next Forcing: Causal World Modeling with Multi-Chunk Prediction (MCP)☆120Aug 16, 2026Updated last week
- PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokens☆18Nov 26, 2025Updated 8 months ago
- Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation☆750Jul 22, 2026Updated last month
- [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations☆202Sep 18, 2025Updated 11 months ago
- ☆39Dec 4, 2023Updated 2 years ago