☆26Jun 6, 2025Updated last year
Alternatives and similar repositories for VLM-Video-Action-Localization
Users that are interested in VLM-Video-Action-Localization are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [WACV 2025] Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection☆17Mar 23, 2025Updated last year
- ☆29Apr 8, 2025Updated last year
- unofficial implementation of physical intelligence pi06☆25Nov 26, 2025Updated 8 months ago
- [NeurIPS 2025 Spotlight] Unleashing Hour-Scale Video Training for Long Video-Language Understanding☆19Jun 24, 2025Updated last year
- Placeholder for code of BSP.☆11Aug 13, 2021Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [CVPR 2026] Official Repository of 'MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos'☆48Jun 22, 2026Updated last month
- Semantic Segmentation in Pytorch☆10Dec 9, 2022Updated 3 years ago
- [ICML 2025] This is the official repository of our paper "What If We Recaption Billions of Web Images with LLaMA-3 ?"☆152Jun 13, 2024Updated 2 years ago
- Public code release for SIGGRAPH 2021 paper: ShapeMOD: Macro Operation Discovery for 3D Shape Programs☆13Sep 8, 2021Updated 4 years ago
- ☆10Nov 10, 2022Updated 3 years ago
- ☆12Sep 29, 2019Updated 6 years ago
- Data release for Step Differences in Instructional Video (CVPR24)☆15Jun 19, 2024Updated 2 years ago
- ☆18Dec 1, 2025Updated 8 months ago
- ☆57Sep 13, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆12Nov 28, 2022Updated 3 years ago
- Lazy Receding Horizon A*☆13Oct 25, 2018Updated 7 years ago
- Official implementation of "Test-Time Zero-Shot Temporal Action Localization", CVPR 2024☆77Sep 11, 2024Updated last year
- Official implementation of ECCV 2024 paper: Take A Step Back: Rethinking the Two Stages in Visual Reasoning☆13Jun 1, 2025Updated last year
- A Unified Framework for Video-Language Understanding☆62Jun 17, 2023Updated 3 years ago
- ☆12Aug 7, 2024Updated 2 years ago
- Code and benchmark of the paper "MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents" (NeurIPS D&B 2025)☆15Oct 13, 2025Updated 9 months ago
- ☆18Oct 20, 2022Updated 3 years ago
- Official Code for ICLR 2023 Paper: A Message Passing Perspective on Learning Dynamics of Contrastive Learning☆11Mar 9, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CVPR 2022] OCSampler: Compressing Videos to One Clip with Single-step Sampling☆17Jun 21, 2022Updated 4 years ago
- ☆17Jun 21, 2024Updated 2 years ago
- ☆11Jul 4, 2024Updated 2 years ago
- ☆13Apr 28, 2019Updated 7 years ago
- This is the official implementation of ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos☆47Nov 5, 2025Updated 9 months ago
- This is a repository contains the implementation of our NeurIPS'24 paper "Temporal Sentence Grounding with Relevance Feedback in Videos"☆13Aug 22, 2025Updated 11 months ago
- [ICCV'25] T2 -VLM: Training-Free Generation of Temporally Consistent Rewards from VLMs☆16Jul 8, 2025Updated last year
- Python implementation of the Tobii Pro Glasses 3 API☆13Sep 12, 2023Updated 2 years ago
- ☆16Apr 14, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆13Mar 18, 2024Updated 2 years ago
- ☆16Apr 8, 2023Updated 3 years ago
- Agentic Keyframe Search for Video Question Answering☆18Jun 30, 2026Updated last month
- [WACV 2026] SceneEdited: A City-Scale Benchmark for 3D HD Map Updating via Image-Guided Change Detection☆19Jul 22, 2026Updated 2 weeks ago
- ☆15Mar 15, 2023Updated 3 years ago
- The Full-Duplex Interaction Track of the ICASSP 2026 Human-like Spoken Dialogue Systems Challenge aims to advance the evaluation of full-…☆38Apr 27, 2026Updated 3 months ago
- Dense classification of the depth images to recognize the body parts.☆19Apr 20, 2020Updated 6 years ago