The official implementation for the paper "Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything".
☆23Nov 5, 2025Updated 8 months ago
Alternatives and similar repositories for Agent-Omni
Users that are interested in Agent-Omni are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repository provides the official implementation of VTBench, a benchmark designed to evaluate the performance of visual tokenizers (V…☆35Jul 30, 2025Updated 11 months ago
- The implementation for paper "UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in …☆17Jul 3, 2025Updated last year
- OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding☆22Apr 9, 2026Updated 3 months ago
- RapidIn: Scalable Influence Estimation for Large Language Models (LLMs). The implementation for paper "Token-wise Influential Training Da…☆22Mar 10, 2026Updated 4 months ago
- Multi-step reasoning MLLM☆25Mar 8, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs☆50Jul 12, 2026Updated last week
- The Source Code for OmniVideoBench @ICLR 2026☆77Feb 12, 2026Updated 5 months ago
- High-Order interactions☆12Jul 5, 2024Updated 2 years ago
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 2 months ago
- LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs☆41Apr 2, 2026Updated 3 months ago
- ☆14Feb 26, 2024Updated 2 years ago
- The DJ Mix Dataset☆20Sep 7, 2022Updated 3 years ago
- Automatically download and crop key information from the arxiv daily paper.☆21Jul 30, 2022Updated 3 years ago
- [ICCV 2025] Official repository of "Mitigating Object Hallucinations via Sentence-Level Early Intervention".☆31Jul 2, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- baseline code for music-crs challenge☆18Jun 23, 2026Updated 3 weeks ago
- code for paper Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering☆14Aug 13, 2024Updated last year
- ☆21Feb 16, 2025Updated last year
- ☆18Aug 19, 2024Updated last year
- FastLongSpeech is a novel framework designed to extend the capabilities of Large Speech-Language Models for efficient long-speech process…☆16Jul 22, 2025Updated last year
- ☆23Feb 4, 2026Updated 5 months ago
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos☆21Jun 20, 2026Updated last month
- Official repo for MMAU-Pro Benchmark☆22Sep 25, 2025Updated 9 months ago
- ☆18Sep 7, 2023Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model☆15Apr 7, 2026Updated 3 months ago
- ☆26Sep 15, 2022Updated 3 years ago
- Enhancing Retrieval and Managing Retrieval: 4-Module Synergy☆23Dec 7, 2024Updated last year
- Official repository for "Boosting Audio Visual Question Answering via Key Semantic-Aware Cues" in ACM MM 2024.☆17Oct 25, 2024Updated last year
- ☆15Jan 12, 2026Updated 6 months ago
- Use 2 lines to empower absolute time awareness for Qwen2.5VL's MRoPE☆29Sep 20, 2025Updated 10 months ago
- [ECCV 2026🔥] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception…☆63Apr 2, 2026Updated 3 months ago
- ☆40May 9, 2024Updated 2 years ago
- [ICLR 2025] Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization☆24Oct 5, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆29Mar 10, 2026Updated 4 months ago
- ☆17Jan 16, 2026Updated 6 months ago
- [NeurIPS 2025] Official Repo of Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration☆126Dec 3, 2025Updated 7 months ago
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering [ACM MM'24]☆10Jul 22, 2024Updated last year
- Distributed In-Memory Trajectory Analytics☆39Jan 22, 2018Updated 8 years ago
- A Multi-Agent Approach Integrating Socratic Guidance for Automated Prompt Optimization☆18Dec 15, 2025Updated 7 months ago
- Multigranularity Contrastive cross-modal collaborative Generation (MCG) model for Video QA☆12Dec 13, 2023Updated 2 years ago