[ECCV 2024] Official PyTorch implementation of "Classification Matters: Improving Video Action Detection with Class-Specific Attention"
☆19Nov 8, 2024Updated last year
Alternatives and similar repositories for class-query-vad
Users that are interested in class-query-vad are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CVPR2022:Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency☆18Aug 10, 2022Updated 4 years ago
- ☆12Aug 7, 2024Updated 2 years ago
- Official Pytorch Implementation of 'BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos'☆36Feb 26, 2025Updated last year
- [NeurIPS 2023] Official implementation of the paper "CAST: Cross-Attention in Space and Time for Video Action Recognition"☆55Dec 28, 2023Updated 2 years ago
- [ICLR 2026] Official implementation of "Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation"☆36Jan 26, 2026Updated 8 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Implementation of Learning Instance-Aware Object Detection Using Determinantal Point Processes [https://arxiv.org/pdf/1805.10765.pdf]☆19Nov 21, 2023Updated 2 years ago
- ☆20Aug 18, 2020Updated 6 years ago
- [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs☆62Feb 2, 2026Updated 7 months ago
- [CVPR 2025] Official Repository of the paper "On the Consistency of Video Large Language Models in Temporal Comprehension"☆16Oct 13, 2025Updated 11 months ago
- We have implemented Track # 1 for ICME 2024: Spatial Action Localization on Chaotic World dataset. Our mAP on the validation set reaches …☆14Nov 11, 2024Updated last year
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos☆26Jun 20, 2026Updated 3 months ago
- Code for "Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning [EMNLP 2025 Findings]"☆18Aug 27, 2025Updated last year
- The Source Code for IF-VidCap @ICLR 2026☆18Oct 22, 2025Updated 11 months ago
- The speaker-labeled information of LRW dataset, which is the outcome of the paper "Speaker-adaptive Lip Reading with User-dependent Paddi…☆10Oct 12, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Official source code for the paper "Tailored Design of Audio-Visual Speech Recognition Models using Branchformers"☆15Feb 24, 2025Updated last year
- Implementation of Deep Elastic Network☆42Nov 24, 2025Updated 10 months ago
- Official Implementation of Video-MA2MBA☆12Dec 3, 2024Updated last year
- Pytorch implementation of "Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal T…☆12Apr 29, 2026Updated 4 months ago
- Implementation of the techniques presented in "Co-occurrence Feature Learning from Skeleton Data for Action Recognition" to recognize two…☆11Jul 22, 2019Updated 7 years ago
- Paper Today I Read☆30Sep 16, 2026Updated last week
- [CVPR 2024] Adapting Short-Term Transformers for Action Detection in Untrimmed Videos☆11Jun 11, 2024Updated 2 years ago
- [ICCV 2025] CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images☆57Jul 15, 2025Updated last year
- ☆15Apr 25, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for "CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally"☆29Aug 3, 2026Updated last month
- Large-Vocabulary Continuous Sign Language Recognition, 2024☆15May 30, 2024Updated 2 years ago
- [ICCV'25] HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics☆37Sep 10, 2025Updated last year
- [WACV 2025] Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection☆17Mar 23, 2025Updated last year
- Implementation of Tsallis Actor Critic method☆61Nov 24, 2025Updated 10 months ago
- [ICASSP 2024] Official code for Slowfast Network for Continuous Sign Language Recognition☆67Jul 4, 2025Updated last year
- FastMIM, official pytorch implementation of our paper "FastMIM: Expediting Masked Image Modeling Pre-training for Vision"(https://arxiv.o…☆39Dec 29, 2022Updated 3 years ago
- Code for "Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations" (CVPR 2024 Oral)☆20Jun 23, 2024Updated 2 years ago
- [CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice☆89Feb 27, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- (CVPR 2026) Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation☆39Feb 28, 2026Updated 6 months ago
- Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation (ACM MM 2024)☆20Mar 17, 2025Updated last year
- ☆11Oct 13, 2024Updated last year
- [CVPR 2026] Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos☆30Mar 25, 2026Updated 6 months ago
- [CVPR 2023] Pytorch Code of MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question Answering☆17Jul 11, 2023Updated 3 years ago
- VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs☆18Feb 3, 2026Updated 7 months ago
- [AAAI 2025] Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos☆40May 27, 2025Updated last year