VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMs
β57Mar 9, 2025Updated last year
Alternatives and similar repositories for VideoNIAH
Users that are interested in VideoNIAH are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π₯π₯MLVU: Multi-task Long Video Understanding Benchmarkβ269Apr 13, 2026Updated 5 months ago
- Official code of *Towards Event-oriented Long Video Understanding*β12Jul 26, 2024Updated 2 years ago
- β43Nov 8, 2024Updated last year
- Official code for the ICLR 2025 paper, "Ada-K Routing: Boosting the Efficiency of MoE-based LLMs"β12Mar 1, 2025Updated last year
- [Neurips 24' D&B] Official Dataloader and Evaluation Scripts for LongVideoBench.β139Jul 27, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ACL 2024 Findings] "TempCompass: Do Video LLMs Really Understand Videos?", Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, β¦β132Apr 4, 2025Updated last year
- The code for "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by VIdeo SpatioTemporal Augmentation" [CVPR2025]β20Feb 27, 2025Updated last year
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoningβ38May 9, 2026Updated 5 months ago
- [COLING 2025π₯] Evolver: Chain-of-Evolution Prompting to Boost Large Multimodal Models for Hateful Meme Detectionβ17Jan 21, 2025Updated last year
- [ICCV 2025] LVBench: An Extreme Long Video Understanding Benchmarkβ155Jul 9, 2025Updated last year
- MR. Video: MapReduce is the Principle for Long Video Understandingβ32Jun 18, 2026Updated 3 months ago
- Long Context Transfer from Language to Visionβ413Mar 18, 2025Updated last year
- [NeurIPS 2024] Needle In A Multimodal Haystack (MM-NIAH): A comprehensive benchmark designed to systematically evaluate the capability ofβ¦β130Nov 25, 2024Updated last year
- Official code for CVPR 2024 paper, "SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models"β16Apr 22, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ECCV 2024π₯] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learners"β156Sep 10, 2024Updated 2 years ago
- β122Dec 30, 2024Updated last year
- β¨β¨The Curse of Multi-Modalities (CMM): Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audioβ55Jul 11, 2025Updated last year
- β18Jul 10, 2024Updated 2 years ago
- The official repo for "VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search" [EMNLP25]β39Feb 1, 2026Updated 8 months ago
- V1: Toward Multimodal Reasoning by Designing Auxiliary Taskβ36Apr 14, 2025Updated last year
- [EMNLP 2025 Main] Official implementation of VRoPE: Rotary Position Embedding for Video Large Language Models.β28Nov 18, 2025Updated 10 months ago
- ChatBridge, an approach to learning a unified multimodal model to interpret, correlate, and reason about various modalities without relyβ¦β55Sep 4, 2023Updated 3 years ago
- β123Feb 4, 2026Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"β79Jan 13, 2026Updated 8 months ago
- Code for paper "Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning"β47Feb 19, 2026Updated 7 months ago
- β¨β¨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysisβ799Dec 8, 2025Updated 10 months ago
- [CVPR'2025] VoCo-LLaMA: This repo is the official implementation of "VoCo-LLaMA: Towards Vision Compression with Large Language Models".β207Jun 18, 2025Updated last year
- [NeurIPS 2024] Artemis: Towards Referential Understanding in Complex Videosβ27Apr 8, 2025Updated last year
- Heterformer: Transformer-based Deep Node Representation Learning on Heterogeneous Text-Rich Networks (KDD 2023)β28Feb 16, 2024Updated 2 years ago
- Official code repo of Video-Browser: Towards Agentic Open-web Video Browsingβ28Jan 19, 2026Updated 8 months ago
- [NeurIPS 2024] Mitigating Object Hallucination via Concentric Causal Attentionβ69Aug 30, 2025Updated last year
- LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architectureβ211Jan 6, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Comprehensive benchmark for video text understandingβ29Jun 4, 2025Updated last year
- FreeVA: Offline MLLM as Training-Free Video Assistantβ70Jun 9, 2024Updated 2 years ago
- π₯π₯First-ever hour scale video understanding modelsβ629Jul 14, 2025Updated last year
- β¨First Open-Source R1-like Video-LLM [2025/02/18]β384Jul 1, 2026Updated 3 months ago
- Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Modelβ282Jun 25, 2024Updated 2 years ago
- (NeurIPS 2024 Spotlight) TOPA: Extend Large Language Models for Video Understanding via Text-Only Pre-Alignmentβ29Sep 27, 2024Updated 2 years ago
- β159Oct 31, 2024Updated last year