Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
☆543Sep 5, 2026Updated this week
Alternatives and similar repositories for Awesome-Multimodal-Modeling
Users that are interested in Awesome-Multimodal-Modeling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- BlogrXiv - AI Research Blog Discovery☆161Updated this week
- Auto-Rubric as Reward: From Implicit Preference to Explicit Generative Criteria☆55Jul 24, 2026Updated last month
- Awesome Multimodal Agent☆112Aug 15, 2026Updated 3 weeks ago
- Unified World Model Inference & Evaluation Infrastructure☆311Jul 22, 2026Updated last month
- ScalingOpt - Optimization Community☆107Jun 1, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights☆33Jan 9, 2026Updated 7 months ago
- Awesome Unified Multimodal Models☆1,317Mar 24, 2026Updated 5 months ago
- Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer☆16Nov 21, 2024Updated last year
- The official code of "Mano: Restriking Manifold Optimization for LLM Training".☆25Jun 1, 2026Updated 3 months ago
- [ICLR 2026] MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding☆23Feb 27, 2026Updated 6 months ago
- [WACV 2026]Official Code of the paper “Equivariant Sampling for Improving Diffusion Model-based Image Restoration“☆19Jan 29, 2026Updated 7 months ago
- 📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.☆366Jan 8, 2026Updated 8 months ago
- [Roadmap] Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling☆134Jun 9, 2026Updated 2 months ago
- SparAlloc: A Simple and Modular Framework for Decoupled Sparsity Allocation in Layerwise Pruning for LLM☆16Jun 5, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Dataflow-MM, multi-media operators for Dataflow. We aim to prepare data for Multimodal Large Language Models.☆48Apr 13, 2026Updated 4 months ago
- WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes☆112Mar 19, 2025Updated last year
- Awesome List for On-Policy Distillation☆852Aug 26, 2026Updated last week
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation☆48Jul 22, 2025Updated last year
- A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts…☆3,400Sep 1, 2026Updated last week
- Code as World: Agentic Discovery of Executable World Representations for Physical Reasoning☆408Aug 31, 2026Updated last week
- 📖 This is a repository for organizing papers, codes and other resources related to unified multimodal models.☆831Oct 10, 2025Updated 10 months ago
- 🌐 Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems☆154Jul 12, 2026Updated last month
- ☆32Apr 29, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of <Model Merging with Functional Dual Anchors>☆46Nov 23, 2025Updated 9 months ago
- OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.☆404Jun 1, 2025Updated last year
- LLMBind: A Unified Modality-Task Integration Framework☆19Jun 16, 2024Updated 2 years ago
- [ECCV 2026] Official implementation of "TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning"☆25Feb 8, 2026Updated 7 months ago
- Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation☆755Jul 22, 2026Updated last month
- A Minimalist, Batteries-included Repository for Advancing World Model Science.☆729Updated this week
- [ICLR 2026] Official implementation (Claude Agent reproduce supported) of paper "mtLoRA: Scalable Multi-Task Low-Rank Model Adaptation" +…☆19Mar 4, 2026Updated 6 months ago
- Official repository for “Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space”☆18Jan 27, 2026Updated 7 months ago
- Implementation for POET and POET-X for LLM pretraining☆41Jun 9, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 本人的科研经验☆13,964Jun 6, 2026Updated 3 months ago
- ☆473Jul 21, 2026Updated last month
- [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s …☆695Feb 27, 2026Updated 6 months ago
- The first multiplayer video world model in Minecraft☆226Mar 3, 2026Updated 6 months ago
- A paper list for spatial reasoning☆780Aug 23, 2026Updated 2 weeks ago
- ☆17Oct 5, 2025Updated 11 months ago
- [ICLR 2026] Offical implementation of "OBS-Diff".☆67Mar 5, 2026Updated 6 months ago