"Object-Region Video Transformers”, Herzig et al., CVPR 2022
☆50Jul 6, 2022Updated 4 years ago
Alternatives and similar repositories for ORViT
Users that are interested in ORViT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code repository for the paper: 'Something-Else: Compositional Action Recognition with Spatial-Temporal Interaction Networks'☆148Aug 25, 2023Updated 3 years ago
- This repository is a fork of https://github.com/joslefaure/HIT customized for the AVA dataset☆17Jun 17, 2023Updated 3 years ago
- N-EPIC-Kitchens: The event-based camera extension of the large-scale EPIC-Kitchens dataset.☆23May 10, 2022Updated 4 years ago
- [ECCV24] VISA: Reasoning Video Object Segmentation via Large Language Model☆22Jul 20, 2024Updated 2 years ago
- A zero-shot captcha solver.☆16Dec 22, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of 3D attention mechanisms based on https://github.com/LeftAttention/Attention-Codebase. Thanks to LeftAttetnion for shari…☆12Feb 22, 2022Updated 4 years ago
- Materials for PyCon 2016 in Portland, Oregon☆10Aug 30, 2015Updated 11 years ago
- Slide and notebook used for my talk on vaex at the Pandas summit 2019 @ Lodnon☆11Jun 13, 2019Updated 7 years ago
- Official TensorFlow code for the paper "DeepWay: a Deep Learning Waypoint Estimator for Global Path Generation".☆11Jun 24, 2022Updated 4 years ago
- Video-Language Alignment via Spatio–Temporal Graph Transformer; ArXiv: https://arxiv.org/abs/2407.11677☆15Jul 24, 2024Updated 2 years ago
- [AAAI'25]: Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP☆23Aug 5, 2025Updated last year
- 🚴♂️ ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction Detection (MM 2020)☆35Jul 2, 2025Updated last year
- Pytorch Implementation of "Object level Visual Reasoning in Videos", F. Baradel, N. Neverova, C. Wolf, J. Mille, G. Mori , ECCV 2018☆169Sep 11, 2018Updated 7 years ago
- [CVPR 2023] STMixer: A One-Stage Sparse Action Detector☆64May 18, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for the paper "Understanding and Evaluating Racial Biases in Image Captioning"☆12Mar 26, 2026Updated 5 months ago
- Official project of DiverseSampling (ACMMM2022 Paper)☆16Feb 25, 2023Updated 3 years ago
- 【ACMMM'2021】DSANet: Dynamic Segment Aggregation Network for Video-Level Representation Learning☆41Jul 7, 2021Updated 5 years ago
- CapsNet implementation in a minimal manner☆11Nov 17, 2017Updated 8 years ago
- Implementation of the paper Video Action Transformer Network☆138Apr 5, 2021Updated 5 years ago
- Is Depth Really Necessary for Salient Object Detection? ACM MM 2020☆22May 30, 2024Updated 2 years ago
- Distributed Training of Bayesian Neural Networks at Scale☆11May 26, 2020Updated 6 years ago
- Official Pytorch Implementation of Relational Self-Attention, NeurIPS 2021☆49Dec 7, 2021Updated 4 years ago
- Official PyTorch implementation of "TDAM: Top-down attention module for CNNs"☆13Oct 29, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICCVW 2023] Interaction-Aware Prompting for Zero-Shot Spatio-Temporal Action Detection☆21Feb 22, 2024Updated 2 years ago
- PyTorch CZSL framework containing GQA, the open-world setting, and the CGE and CompCos methods.☆129Oct 29, 2025Updated 10 months ago
- DL4CV book☆10Sep 18, 2018Updated 7 years ago
- A Simulator for Traffic Intersection based on Crossroads technique☆10Dec 4, 2019Updated 6 years ago
- EPIC-Kitchens-100 Action Recognition baselines: TSN, TRN, TSM☆33Mar 15, 2022Updated 4 years ago
- Official implementation for "GLASS: Global to Local Attention for Scene-Text Spotting" (ECCV'22)☆102Jun 28, 2024Updated 2 years ago
- [CVPR 2024 Challenge] 1st Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation☆31Oct 18, 2024Updated last year
- Multi-head Recurrent Layer Attention for Vision Network☆23Mar 2, 2023Updated 3 years ago
- A simple tkinter GUI for illustrating DFS and BFS.☆10Jun 26, 2020Updated 6 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆107Dec 23, 2022Updated 3 years ago
- A simple way to transport dynamic data over ROS comms☆17Aug 15, 2026Updated 2 weeks ago
- Interaction Compass: Multi-Label Zero-Shot Learning of Human-Object Interactions via Spatial Relations @ ICCV21☆13Jul 15, 2022Updated 4 years ago
- A Rideshare Simulation built in C++, using OpenStreetMap data☆14Oct 24, 2021Updated 4 years ago
- The code repository for "Cross-Modal and Hierarchical Modeling of Video and Text" in PyTorch☆16Apr 22, 2019Updated 7 years ago
- A Probabilistic Programming Language in 70 lines of Python. Code for the blog post https://mrandri19.github.io/2022/01/12/a-PPL-in-70-lin…☆19Feb 10, 2022Updated 4 years ago
- Official Implementation of our WACV2023 paper: “Holistic Interaction Transformer Network for Action Detection”☆72Jan 9, 2025Updated last year