Comprehensive benchmark for video text understanding
☆29Jun 4, 2025Updated last year
Alternatives and similar repositories for VidText
Users that are interested in VidText are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 🔥🔥[NeurIPS2025]Exploring and mitigating semantic hallucinations in scene text perception and reasoning☆30Dec 11, 2025Updated 9 months ago
- DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning☆20Updated this week
- Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"☆20Jan 18, 2026Updated 8 months ago
- Official code repo of Video-Browser: Towards Agentic Open-web Video Browsing☆28Jan 19, 2026Updated 8 months ago
- (AAAI2026) Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning (CAL)☆32Jan 1, 2026Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMs☆57Mar 9, 2025Updated last year
- ☆29Aug 9, 2025Updated last year
- EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams [CVPR'24]☆33Jul 23, 2025Updated last year
- ☆32Jul 29, 2024Updated 2 years ago
- 🔥🔥MLVU: Multi-task Long Video Understanding Benchmark☆269Apr 13, 2026Updated 5 months ago
- Official InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows☆20Nov 4, 2025Updated 10 months ago
- Multiple-Person Multi-Camera Tracker☆13Feb 17, 2017Updated 9 years ago
- ☆11Nov 27, 2025Updated 10 months ago
- ☆15Apr 25, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 【AAAI 2022】Temporal Action Proposal Generation with Background Constraint☆17May 13, 2022Updated 4 years ago
- [IEEE TMM'25] Scene-Text Grounding for Text-Based Video Question Answering☆17Feb 16, 2026Updated 7 months ago
- [ACCV 2024 (Oral, Best Application Paper)] Official Implementation of NT-VOT211: A Large-Scale Benchmark for Night-time Visual Object Tra…☆16Dec 30, 2025Updated 9 months ago
- 🔥🔥First-ever hour scale video understanding models☆630Jul 14, 2025Updated last year
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 5 months ago
- [2024-NeurIPS] TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control☆107Mar 16, 2025Updated last year
- The repo for code, that hasn't been published yet☆15Updated this week
- Repo for paper "T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs"☆48Sep 3, 2025Updated last year
- We introduce DreamPRM-1.5, an instance-reweighted framework that adaptively adjusts the importance of each training example via bi-level …☆16Nov 13, 2025Updated 10 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ECCV2024] ModTr: Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledge☆20Nov 28, 2024Updated last year
- ☆13Jun 1, 2023Updated 3 years ago
- Fleming-VL: Towards Universal Medical Visual Understanding with Multimodal LLMs☆18Nov 6, 2025Updated 10 months ago
- An official Project related to Paper "Perceiving Ambiguity and Semantics without Recognition: An Efficient and Effective Ambiguous Scene …☆22Dec 3, 2023Updated 2 years ago
- Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding☆40Mar 16, 2025Updated last year
- Turn every moment into momentum☆22Jun 1, 2026Updated 4 months ago