[CVPR'26 Demo] Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
☆154Apr 13, 2026Updated 3 months ago
Alternatives and similar repositories for Mobile-O
Users that are interested in Mobile-O are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VideoMathQA is a benchmark designed to evaluate mathematical reasoning in real-world educational videos☆24May 7, 2026Updated 2 months ago
- WorldCache: Content-Aware Caching for Accelerated Video World Models☆21Jun 28, 2026Updated 3 weeks ago
- [CVPR2025] Official Implementations "One-Way Ticket : Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models"☆29Mar 16, 2026Updated 4 months ago
- Self Evolving Large Multimodal Models with Continuous Rewards☆25Jun 9, 2026Updated last month
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations☆22Jun 17, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models☆19Jan 21, 2026Updated 6 months ago
- ☆15Jun 2, 2025Updated last year
- [ArXiv 2025] MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices☆87May 20, 2026Updated 2 months ago
- [ICLR 2026] Code for Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model☆30Mar 1, 2026Updated 4 months ago
- [WACV 2025] Efficient Video Object Segmentation via Modulated Cross-Attention Memory☆61Feb 28, 2025Updated last year
- [ICML2026] Official Implementations "FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models"☆27Jul 9, 2026Updated last week
- OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning☆29May 24, 2025Updated last year
- ICLR 2026: Agent-X Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks☆43Apr 28, 2026Updated 2 months ago
- Mobile-VideoGPT: Fast and Accurate Video Understanding Language Model☆142Aug 6, 2025Updated 11 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator☆33Apr 15, 2026Updated 3 months ago
- [ICLR 2026 🔥] Dr.LLM: Dynamic Layer Routing in LLMs☆56Apr 24, 2026Updated 2 months ago
- [CVPR -2025] GroupMamba: Parameter-Efficient and Accurate Group Visual State Space Model☆142Mar 22, 2025Updated last year
- InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image edit…☆291Mar 21, 2026Updated 4 months ago
- AAAI2026 X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning☆97Nov 21, 2025Updated 8 months ago
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos☆21Jun 20, 2026Updated last month
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"☆56Jul 5, 2025Updated last year
- Nitro-E is a family of text-to-image diffusion models focused on highly efficient training.☆125Jun 4, 2026Updated last month
- [ECCV 2026] 🔥 Official impl. of "DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing".☆731Jun 12, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆55Feb 9, 2026Updated 5 months ago
- [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers☆511Dec 6, 2025Updated 7 months ago
- ☆26Nov 25, 2025Updated 7 months ago
- [Arxiv 2025] ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions☆45Jun 11, 2025Updated last year
- [ICCV 2025] Factorized Learning for Temporally Grounded Video-Language Models☆24Apr 18, 2026Updated 3 months ago
- [CVPR 2025] Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models☆365Nov 24, 2025Updated 7 months ago
- Official Implementations "Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference" for DiT (NeurIPS'24)☆15Aug 3, 2025Updated 11 months ago
- Official Repo for Paper <EditMGT Unleashing the Potential of Masked Generative Transformer in Image Editing>☆79Dec 20, 2025Updated 7 months ago
- SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation (CVPR 2024)☆72Jun 24, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow Distillation☆128May 20, 2026Updated 2 months ago
- [ACM MM25] LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models☆24Mar 29, 2025Updated last year
- ☆20Jul 14, 2024Updated 2 years ago
- Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders☆255Feb 13, 2026Updated 5 months ago
- NextFlow🚀: Unified Sequential Modeling Activates Multimodal Understanding and Generation☆331Jan 9, 2026Updated 6 months ago
- Offical repo for "OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution"☆118Mar 11, 2026Updated 4 months ago
- Official Implementation for "Transferring Unconditional to Conditional GANs with Hyper-Modulation" CVPRW 22 https://arxiv.org/abs/2112.02…☆13Jun 28, 2022Updated 4 years ago