Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
☆369Aug 5, 2026Updated last week
Alternatives and similar repositories for Video-MME-v2
Users that are interested in Video-MME-v2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆30May 21, 2026Updated 2 months ago
- VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding☆58May 1, 2026Updated 3 months ago
- [CVPR 2026 Highlight] PersonaVLM: Long-Term Personalized Multimodal LLMs☆115Apr 16, 2026Updated 3 months ago
- ✨✨[ICML 2026] Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion☆153Mar 12, 2026Updated 5 months ago
- ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis☆789Dec 8, 2025Updated 8 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- The Next Step Forward in Multimodal LLM Alignment☆199May 1, 2025Updated last year
- ✨✨Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy☆306May 14, 2025Updated last year
- EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory☆53Jul 24, 2026Updated 3 weeks ago
- Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation☆32Mar 28, 2025Updated last year
- [ECCV'26] GRADE: Grounded Reasoning Assessment for Discipline-informed Editing☆29Apr 23, 2026Updated 3 months ago
- ✨✨[AAAI 2026] This is the official implementation of our paper "QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Vi…☆79Apr 28, 2025Updated last year
- The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.☆164Oct 28, 2025Updated 9 months ago
- A list of awesome papers on compression and acceleration of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs).☆17May 12, 2026Updated 3 months ago
- PhotoFlow: Agentic 3D Virtual Photography Missions☆38May 27, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆20Jun 2, 2026Updated 2 months ago
- [ECCV'26] Code repo for "EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation"☆22Jun 18, 2026Updated last month
- RISE-Video: Can Video Generators Decode Implicit World Rules?☆28Mar 26, 2026Updated 4 months ago
- ✨✨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models☆43Apr 10, 2025Updated last year
- ✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models☆649Dec 23, 2024Updated last year
- ☆38Jul 9, 2024Updated 2 years ago
- ☆32Jul 29, 2024Updated 2 years ago
- ☆27Feb 3, 2026Updated 6 months ago
- Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition☆42Mar 3, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆395Jun 20, 2026Updated last month
- [TPAMI 2021] DVG-Face: Dual Variational Generation for Heterogeneous Face Recognition☆76Nov 13, 2023Updated 2 years ago
- SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation☆31Jul 9, 2026Updated last month
- Repo for paper "T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs"☆48Sep 3, 2025Updated 11 months ago
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆31Apr 2, 2026Updated 4 months ago
- ☆36Apr 13, 2026Updated 4 months ago
- ☆23Apr 11, 2026Updated 4 months ago
- ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction