Comprehensive benchmark for video text understanding
β29Jun 4, 2025Updated last year
Alternatives and similar repositories for VidText
Users that are interested in VidText are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π₯π₯[NeurIPS2025]Exploring and mitigating semantic hallucinations in scene text perception and reasoningβ30Dec 11, 2025Updated 9 months ago
- Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"β20Jan 18, 2026Updated 7 months ago
- Official code repo of Video-Browser: Towards Agentic Open-web Video Browsingβ28Jan 19, 2026Updated 7 months ago
- (AAAI2026) Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning (CAL)β32Jan 1, 2026Updated 8 months ago
- Official Code for TPAMI 2024 paper "EvHandPose: Event-based 3D Hand Pose Estimation with Sparse Supervision"β19Dec 4, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMsβ57Mar 9, 2025Updated last year
- π₯π₯A Family of Multi-Sensor, Multi-Granularity Vision-Language Models for Earth Observation Understandingβ142Jun 15, 2026Updated 2 months ago
- Loomis Painter: Reconstructing the painting processβ55Nov 24, 2025Updated 9 months ago
- β28Aug 9, 2025Updated last year
- γ2024 ECAIγFirst Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blendingβ14Jun 16, 2025Updated last year
- EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams [CVPR'24]β33Jul 23, 2025Updated last year
- β32Jul 29, 2024Updated 2 years ago
- π₯π₯MLVU: Multi-task Long Video Understanding Benchmarkβ268Apr 13, 2026Updated 5 months ago
- Official InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Showsβ20Nov 4, 2025Updated 10 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Add YOLOv3_tiny and data augment(clip, brighten, change saturation)β14Jan 14, 2021Updated 5 years ago
- Optocal Character Recognition (OCR / HTR) using Transformersβ11Aug 20, 2022Updated 4 years ago
- Multiple-Person Multi-Camera Trackerβ13Feb 17, 2017Updated 9 years ago
- β11Nov 27, 2025Updated 9 months ago
- β15Apr 25, 2025Updated last year
- β13May 17, 2025Updated last year
- [IEEE TMM'25] Scene-Text Grounding for Text-Based Video Question Answeringβ17Feb 16, 2026Updated 6 months ago
- π₯π₯First-ever hour scale video understanding modelsβ629Jul 14, 2025Updated last year
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Miningβ55Apr 22, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [2024-NeurIPS] TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Controlβ107Mar 16, 2025Updated last year
- The repo for code, that hasn't been published yetβ14May 14, 2025Updated last year
- We introduce DreamPRM-1.5, an instance-reweighted framework that adaptively adjusts the importance of each training example via bi-level β¦β16Nov 13, 2025Updated 10 months ago
- [NeurIPS 2025] AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understandingβ15Aug 10, 2026Updated last month
- [ECCV2024] ModTr: Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledgeβ20Nov 28, 2024Updated last year
- β13Jun 1, 2023Updated 3 years ago
- Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understandingβ40Mar 16, 2025Updated last year
- Turn every moment into momentumβ22Jun 1, 2026Updated 3 months ago
- β57Mar 19, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Rui Qian, Xin Lai, Xirong Li: BADet: Boundary-Aware 3D Object Detection from Point Clouds (Pattern Recognition 2022: IF=8.518)β13Feb 12, 2026Updated 7 months ago
- η«―ε°η«―ηδΈζεΊζ―ζεθ―ε«γβ12Jun 27, 2022Updated 4 years ago
- Responsible Robotic Manipulationβ17Aug 31, 2025Updated last year
- β16Jan 27, 2026Updated 7 months ago
- β16May 1, 2026Updated 4 months ago
- β54Oct 20, 2025Updated 10 months ago
- β15Dec 15, 2025Updated 8 months ago