Comprehensive benchmark for video text understanding
β29Jun 4, 2025Updated last year
Alternatives and similar repositories for VidText
Users that are interested in VidText are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π₯π₯[NeurIPS2025]Exploring and mitigating semantic hallucinations in scene text perception and reasoningβ30Dec 11, 2025Updated 8 months ago
- DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoningβ18Jun 14, 2026Updated 2 months ago
- Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"β20Jan 18, 2026Updated 7 months ago
- Official code repo of Video-Browser: Towards Agentic Open-web Video Browsingβ28Jan 19, 2026Updated 7 months ago
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMsβ57Mar 9, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Loomis Painter: Reconstructing the painting processβ55Nov 24, 2025Updated 9 months ago
- β28Aug 9, 2025Updated last year
- γ2024 ECAIγFirst Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blendingβ14Jun 16, 2025Updated last year
- EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams [CVPR'24]β33Jul 23, 2025Updated last year
- π₯π₯MLVU: Multi-task Long Video Understanding Benchmarkβ268Apr 13, 2026Updated 4 months ago
- β32Jul 29, 2024Updated 2 years ago
- Official InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Showsβ20Nov 4, 2025Updated 9 months ago
- Add YOLOv3_tiny and data augment(clip, brighten, change saturation)β14Jan 14, 2021Updated 5 years ago
- This demo demonstrates the AI capabilities of the mcxn947. It displays the image captured by the camera on the LCD screen and performs faβ¦β14May 18, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Multiple-Person Multi-Camera Trackerβ13Feb 17, 2017Updated 9 years ago
- β11Nov 27, 2025Updated 8 months ago
- β15Apr 25, 2025Updated last year
- β13May 17, 2025Updated last year
- [IEEE TMM'25] Scene-Text Grounding for Text-Based Video Question Answeringβ17Feb 16, 2026Updated 6 months ago
- [ACCV 2024 (Oral, Best Application Paper)] Official Implementation of NT-VOT211: A Large-Scale Benchmark for Night-time Visual Object Traβ¦β16Dec 30, 2025Updated 7 months ago
- π₯π₯First-ever hour scale video understanding modelsβ626Jul 14, 2025Updated last year
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Miningβ55Apr 22, 2026Updated 4 months ago
- [2024-NeurIPS] TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Controlβ107Mar 16, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Repo for paper "T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs"β48Sep 3, 2025Updated 11 months ago
- [ECCV2024] ModTr: Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledgeβ20Nov 28, 2024Updated last year
- [NeurIPS 2025] AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understandingβ15Aug 10, 2026Updated 2 weeks ago
- [CVPR'26 Highlight] MemCoach: Steering-based MLLM for Actionable Image Memorability Feedbackβ44Jul 24, 2026Updated last month
- Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understandingβ40Mar 16, 2025Updated last year
- β57Mar 19, 2025Updated last year
- Rui Qian, Xin Lai, Xirong Li: BADet: Boundary-Aware 3D Object Detection from Point Clouds (Pattern Recognition 2022: IF=8.518)β13Feb 12, 2026Updated 6 months ago
- η«―ε°η«―ηδΈζεΊζ―ζεθ―ε«γβ12Jun 27, 2022Updated 4 years ago
- The official project of paper "Visual Text Processing: A Comprehensive Review and Unified Evaluation""β104Oct 20, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- β16Jan 27, 2026Updated 6 months ago
- β16May 1, 2026Updated 3 months ago
- β15Dec 15, 2025Updated 8 months ago
- β15Jun 22, 2026Updated 2 months ago
- A new video text spotting framework with Transformerβ82May 23, 2022Updated 4 years ago
- F2F-AP: Flow-to-Future Asynchronous Policy for Real-time Dynamic Manipulationβ16Apr 7, 2026Updated 4 months ago
- Use 2 lines to empower absolute time awareness for Qwen2.5VL's MRoPEβ29Sep 20, 2025Updated 11 months ago