Awesome Visual Agent
☆19Jul 1, 2026Updated 3 weeks ago
Alternatives and similar repositories for Awesome-Multimodal-Agent
Users that are interested in Awesome-Multimodal-Agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Auto-Rubric as Reward: From Implicit Preference to Explicit Generative Criteria☆50Updated this week
- BlogrXiv - AI Research Blog Discovery☆125Updated this week
- ScalingOpt - Optimization Community☆104Jun 1, 2026Updated last month
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights☆32Jan 9, 2026Updated 6 months ago
- Unified World Model Inference & Evaluation Infrastructure☆262Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]☆505Updated this week
- Authors implementation of "Flowception Temporally Expansive Flow Matching for Video Generation".☆21May 9, 2026Updated 2 months ago
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward☆60Nov 27, 2025Updated 7 months ago
- ☆13Apr 28, 2025Updated last year
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding (CVPR 2025 Oral)☆42Nov 28, 2025Updated 7 months ago
- ☆15Jul 25, 2024Updated last year
- A tiny paper rating web☆41Mar 19, 2025Updated last year
- ☆17Jan 28, 2024Updated 2 years ago
- Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory☆115Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆20May 28, 2025Updated last year
- vulkan graphic engine for learning☆12Dec 21, 2024Updated last year
- [ICLR26] Understanding VS. Generation: Navigating Optimization Dilemma in Multimodal Models☆25May 6, 2026Updated 2 months ago
- Official eval code for ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation☆26Dec 12, 2025Updated 7 months ago
- WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes☆111Mar 19, 2025Updated last year
- [ICCV 2023] Data-Free Class-Incremental Hand Gesture Recognition☆17Sep 21, 2023Updated 2 years ago
- nerual network enhanced global ilumination with Unity☆11Apr 12, 2024Updated 2 years ago
- Some thoughts about writing scientific papers☆23Nov 8, 2024Updated last year
- 📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.☆365Jan 8, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆31Mar 29, 2026Updated 3 months ago
- ☆12Apr 25, 2025Updated last year
- Collection of recent methods on 3D Scene Generation from Text Description.☆17Mar 3, 2025Updated last year
- [ICML 2026] TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models☆25Mar 24, 2026Updated 4 months ago
- Streaming Video Instruction Tuning☆82Feb 25, 2026Updated 5 months ago
- Offical Repository for Paper: DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation☆17Dec 7, 2025Updated 7 months ago
- SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation. A typed knowledge graph unifies data synthe…☆20Jul 8, 2026Updated 2 weeks ago
- [ICML 2026🔥] WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation☆212Jun 26, 2026Updated 3 weeks ago
- [CGF] A curated list of papers on feed-forward 3D reconstruction and novel view synthesis.☆17Mar 14, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Aug 7, 2025Updated 11 months ago
- A comprehensive benchmark suite for multi-view generation models☆21Jul 14, 2025Updated last year
- [ICML 2026] ScalingAR: Scaling Confidence for Autoregressive Image Generation☆22May 5, 2026Updated 2 months ago
- The official code base for paper AAAI2026 <MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language …☆22Feb 2, 2026Updated 5 months ago
- A Heterogeneous Graph Transformer (HGT)-based model for protein function prediction using biological knowledge graphs and protein languag…☆19Jan 27, 2026Updated 5 months ago
- An Empirical Study of GPT-4o Image Generation Capabilities☆29Apr 16, 2025Updated last year
- [ECCV 2026🔥] SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models☆93Nov 26, 2025Updated 7 months ago