☆23Dec 5, 2023Updated 2 years ago
Alternatives and similar repositories for QuerYD_downloader
Users that are interested in QuerYD_downloader are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A collection of videos annotated with timelines where each video is divided into segments, and each segment is labelled with a short free…☆30Jan 15, 2022Updated 4 years ago
- The official implementation of the paper **LVChat: Facilitating Long Video Comprehension**☆14Apr 15, 2024Updated 2 years ago
- Code for Deep Multimodal Clustering for Unsupervised Audiovisual Learning (CVPR2019)☆15May 27, 2020Updated 6 years ago
- Repository for MarioQA: Answering Questions by Watching Gameplay Videos in ICCV 2017☆10Oct 28, 2025Updated 11 months ago
- ☆26Jun 25, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for the C2KD paper (ICASSP 2023)☆20May 15, 2023Updated 3 years ago
- An Implementation of "Small steps and giant leaps: Minimal Newton solvers for Deep Learning" In pytorch☆21Jul 16, 2018Updated 8 years ago
- When can you tell whether an image has been cropped or not?☆29Sep 19, 2021Updated 5 years ago
- Code for the AVLnet (Interspeech 2021) and Cascaded Multilingual (Interspeech 2021) papers.☆54Mar 30, 2022Updated 4 years ago
- Unsupervised phone and word segmentation using dynamic programming on self-supervised VQ features.☆39May 5, 2026Updated 5 months ago
- A framework for local feature evaluation for Python and MATLAB.☆13Jul 6, 2023Updated 3 years ago
- Cross-Self KV Cache Pruning for Efficient Vision-Language Inference☆10Dec 15, 2024Updated last year
- Audio Visual Instance Discrimination with Cross-Modal Agreement☆133Aug 13, 2021Updated 5 years ago
- Seeing Wake Words: Audio-visual Keyword Spotting☆67Sep 16, 2020Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR25] Official Implementation of CAV-MAE Sync☆34Apr 5, 2026Updated 6 months ago
- Code for the paper: Audio-Visual Model Distillation Using Acoustic Images☆21Mar 24, 2023Updated 3 years ago
- Shapley values for assessing the importance of each frame in a video☆17Mar 1, 2021Updated 5 years ago
- S3D Text-Video model trained on HowTo100M using MIL-NCE☆200Jul 3, 2020Updated 6 years ago
- DO with Terraform and Ansible☆11Jun 5, 2018Updated 8 years ago
- Official github repo for ICCV2023 paper 'Multi-event Video-Text Retrieval'☆20Feb 16, 2024Updated 2 years ago
- ICCV 2021☆34May 11, 2022Updated 4 years ago
- 12-in-1: Multi-Task Vision and Language Representation Learning Web Demo☆35Dec 8, 2022Updated 3 years ago
- Least-squares Reverse Time Migration using 1D scalar wave equation. Very simple and for demonstration purposes only.☆12Sep 4, 2017Updated 9 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implements attacks and defenses for machine learning systems☆13May 7, 2017Updated 9 years ago
- A framework for building speech-enabled websites.☆10Jul 10, 2015Updated 11 years ago
- [ECCV 2020] PyTorch code of MMT (a multimodal transformer captioning model) on TVCaption dataset☆91Sep 6, 2023Updated 3 years ago
- ☆12May 22, 2022Updated 4 years ago
- Code for "Audio Retrieval with Natural Language Queries: A Benchmark Study", Transactions on Multimedia 2022☆53Jul 16, 2025Updated last year
- Cross Modal Retrieval with Querybank Normalisation☆57Nov 21, 2023Updated 2 years ago
- A collection of basic text processing modules focused on Gujarati☆10Oct 24, 2017Updated 8 years ago
- Official InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows☆20Nov 4, 2025Updated 11 months ago
- Research code for "Towards multi-task learning of speech and speaker recognition" at https://arxiv.org/pdf/2302.12773.pdf☆12Dec 2, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆18Dec 10, 2025Updated 9 months ago
- X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization, CVPR 2024☆11Nov 7, 2024Updated last year
- An extension of thu-spmi/CAT which contains a full-fledged implementation of CTC-CRF for Tensorflow.☆12Jul 5, 2021Updated 5 years ago
- Code release for paper: The Boombox: Visual Reconstruction from Acoustic Vibrations☆15May 18, 2021Updated 5 years ago
- Code associated with the paper: CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition.☆17May 16, 2025Updated last year
- ☆12May 26, 2023Updated 3 years ago
- Feedforward implementation of Lightweight Probabilistic Deep Networks for Keras and Tensorflow☆14Jul 1, 2019Updated 7 years ago