[ECCV 2024] Official PyTorch implementation of "Classification Matters: Improving Video Action Detection with Class-Specific Attention"
☆19Nov 8, 2024Updated last year
Alternatives and similar repositories for class-query-vad
Users that are interested in class-query-vad are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CVPR2022:Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency☆18Aug 10, 2022Updated 4 years ago
- ☆12Aug 7, 2024Updated 2 years ago
- Official Pytorch Implementation of 'BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos'☆36Feb 26, 2025Updated last year
- [NeurIPS 2023] Official implementation of the paper "CAST: Cross-Attention in Space and Time for Video Action Recognition"☆55Dec 28, 2023Updated 2 years ago
- [ICLR 2026] Official implementation of "Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation"☆36Jan 26, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of Learning Instance-Aware Object Detection Using Determinantal Point Processes [https://arxiv.org/pdf/1805.10765.pdf]☆19Nov 21, 2023Updated 2 years ago
- Deep learning tutorials using tensorflow☆22Oct 11, 2019Updated 6 years ago
- [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs☆62Feb 2, 2026Updated 7 months ago
- [CVPR 2025] Official Repository of the paper "On the Consistency of Video Large Language Models in Temporal Comprehension"☆16Oct 13, 2025Updated 10 months ago
- We have implemented Track # 1 for ICME 2024: Spatial Action Localization on Chaotic World dataset. Our mAP on the validation set reaches …☆14Nov 11, 2024Updated last year
- Model predictive control under STL constraints☆33Nov 24, 2025Updated 9 months ago
- ☆29Sep 3, 2019Updated 7 years ago
- ☆33Nov 24, 2025Updated 9 months ago
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos☆23Jun 20, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Code for "Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning [EMNLP 2025 Findings]"☆18Aug 27, 2025Updated last year
- The Source Code for IF-VidCap @ICLR 2026☆18Oct 22, 2025Updated 10 months ago
- The speaker-labeled information of LRW dataset, which is the outcome of the paper "Speaker-adaptive Lip Reading with User-dependent Paddi…☆10Oct 12, 2023Updated 2 years ago
- Official source code for the paper "Tailored Design of Audio-Visual Speech Recognition Models using Branchformers"☆15Feb 24, 2025Updated last year
- Implementation of Deep Elastic Network☆42Nov 24, 2025Updated 9 months ago
- Pose refinement with differentiable rendering☆10Dec 27, 2020Updated 5 years ago
- Official Implementation of Video-MA2MBA☆12Dec 3, 2024Updated last year
- Implementation of Unsupervised 3D Reconstruction Network☆46Nov 24, 2025Updated 9 months ago
- Pytorch implementation of "Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal T…☆12Apr 29, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Implementation of the techniques presented in "Co-occurrence Feature Learning from Skeleton Data for Action Recognition" to recognize two…☆11Jul 22, 2019Updated 7 years ago
- Paper Today I Read☆30Updated this week
- [CVPR 2024] Adapting Short-Term Transformers for Action Detection in Untrimmed Videos☆11Jun 11, 2024Updated 2 years ago
- [ICCV 2025] CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images☆57Jul 15, 2025Updated last year
- ☆15Apr 25, 2025Updated last year
- This is the official resources for ECCV 2022 paper "Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation From M…☆18Jun 15, 2023Updated 3 years ago
- Code for "CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally"☆29Aug 3, 2026Updated last month
- Large-Vocabulary Continuous Sign Language Recognition, 2024☆15May 30, 2024Updated 2 years ago
- [ICCV'25] HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics☆37Sep 10, 2025Updated 11 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [WACV 2025] Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection☆17Mar 23, 2025Updated last year
- Crawler for annual (biennial) AI conference papers☆10Dec 30, 2024Updated last year
- [ICASSP 2024] Official code for Slowfast Network for Continuous Sign Language Recognition☆67Jul 4, 2025Updated last year
- A fix of https://pypi.python.org/pypi/mathutils☆17Nov 16, 2016Updated 9 years ago
- Code for "Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations" (CVPR 2024 Oral)☆20Jun 23, 2024Updated 2 years ago
- [CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice☆89Feb 27, 2026Updated 6 months ago
- (CVPR 2026) Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation☆39Feb 28, 2026Updated 6 months ago