Official implementation of "Think, Then Verify: A Hypothesis–Verification Multi-Agent Framework for Long Video Understanding(CVPR'2026)"
☆29Aug 10, 2026Updated last month
Alternatives and similar repositories for VideoHV-Agent
Users that are interested in VideoHV-Agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆23Feb 21, 2026Updated 7 months ago
- [ACL-26 (main)] From Verbatim to Gist Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video A…☆43Apr 19, 2026Updated 5 months ago
- ☆14Jan 9, 2025Updated last year
- ☆11Apr 26, 2024Updated 2 years ago
- The official implementation of VLPL: Vision Language Pseudo Label for Multi-label Learning with Single Positive Labels☆18Jul 16, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [MICCAI2025] A latent motion profiling method for unsupervised cardiac phase detection.☆17Dec 15, 2025Updated 9 months ago
- This repository is the official Pytorch implementation of Balanced Product of Calibrated Experts for Long-Tailed Recognition (CVPR 2023).☆16Mar 13, 2025Updated last year
- [CVPR 2026] VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking☆72Mar 23, 2026Updated 6 months ago
- Official Code of "Random Parameter Pruning Attack (Accepeted by CVPR26)"☆17Feb 26, 2026Updated 7 months ago
- Visual Speech Recongnition☆22Dec 24, 2024Updated last year
- ☆21May 15, 2026Updated 4 months ago
- code of cvpr26 paper Symphony☆17Apr 7, 2026Updated 5 months ago
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆20Jun 2, 2026Updated 4 months ago
- [ECCV 2026] StAR: Segment Anything Reasoner☆26Apr 2, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation of "AgentRVOS: Reasoning Over Object Tracks for Zero-Shot Referring Video Object Segmentation".☆23Mar 25, 2026Updated 6 months ago
- Code for EMNLP25 paper "Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning"☆24Feb 18, 2026Updated 7 months ago
- [ACMMM'25] Referring Expression Instance Retrieval and A Strong End-to-End Baseline☆19Apr 7, 2026Updated 5 months ago
- ☆21Jan 17, 2025Updated last year
- [CVPR 2026] WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval☆26Jun 17, 2026Updated 3 months ago
- ☆26Apr 7, 2025Updated last year
- Code implementation of paper "MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval (AAAI2025)"☆26Feb 2, 2025Updated last year
- [CVPR2022] Official Implementation of the paper 'Learning Where to Learn in Cross-View Self-Supervised Learning'☆29Oct 12, 2022Updated 3 years ago
- [ICLR 2026 Oral] Official Implementation of the paper "MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interactio…☆24Jul 2, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of Poincare Embedding in PyTorch☆13Jul 27, 2017Updated 9 years ago
- [ICML 2026] HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling☆31May 2, 2026Updated 5 months ago
- ☆21Apr 21, 2026Updated 5 months ago
- Cross-State Transition Attention Transformer for improved robotic manipulation with better temporal modeling; https://arxiv.org/abs/2510.…☆19Mar 8, 2026Updated 6 months ago
- ☆14Feb 26, 2024Updated 2 years ago
- ☆12Jul 14, 2022Updated 4 years ago
- EVA: Efficient Reinforcement Learning for End-to-End Video Agent☆26May 6, 2026Updated 4 months ago
- The code repository for "Cross-Modal and Hierarchical Modeling of Video and Text" in PyTorch☆20Apr 26, 2020Updated 6 years ago
- ☆20Feb 23, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding☆26Apr 29, 2026Updated 5 months ago
- Code release for the paper "Progress-Aware Video Frame Captioning" (CVPR 2025)☆27Jul 16, 2025Updated last year
- ☆12Feb 7, 2018Updated 8 years ago
- CPLEX code of the E-VRPTW☆17Apr 8, 2024Updated 2 years ago
- [arXiv'26] From Clinical Intent to Clinical Model: An Autonomous Coding-Agent Framework for Clinician-driven AI Development☆27Jul 1, 2026Updated 3 months ago
- ☆26Jan 29, 2026Updated 8 months ago
- ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos☆25May 21, 2025Updated last year