repo for active speaker detection for media videos.
☆31Nov 19, 2023Updated 2 years ago
Alternatives and similar repositories for movie-asd
Users that are interested in movie-asd are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repository is a repository for the paper, "Irgun: Improved residue based gradual up-scaling network for single image super resolutio…☆16Aug 26, 2020Updated 5 years ago
- The repository for IEEE CVPR 2023 (A Light Weight Model for Active Speaker Detection)☆186Mar 23, 2025Updated last year
- Official implementation of TalkNCE (ICASSP 2024).☆18Apr 30, 2025Updated last year
- [Interspeech 2026] Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness☆24Jun 25, 2026Updated last month
- PyTorch implementation of USR 2.0 (ICLR 2026)☆15Apr 3, 2026Updated 4 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [INTERSPEECH 2026] Official code for "Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech"☆21Aug 9, 2026Updated last week
- Official implementation of Transpotter, published in BMVC 2021☆16Aug 6, 2022Updated 4 years ago
- Code implementation for our ICPR, 2020 paper titled "Improving Word Recognition using Multiple Hypotheses and Deep Embeddings"☆21May 21, 2021Updated 5 years ago
- Accepted by TMM 2022☆19Aug 18, 2022Updated 4 years ago
- ☆10Dec 22, 2023Updated 2 years ago
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection (ECCV 2022)☆67Oct 29, 2023Updated 2 years ago
- Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question Answering☆11Feb 16, 2023Updated 3 years ago
- Weakly Supervised CRNN System for Sound Event Detection With Large-scale Unlabeled In-domain Data☆11Oct 31, 2018Updated 7 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official code for the paper "GestSync: Determining who is speaking without a talking head" published at BMVC 2023☆48Sep 1, 2024Updated last year
- A Conversational Speech Generation Model☆14Mar 16, 2025Updated last year
- Codebase for "Channel selection using Gumbel Softmax"☆19Jan 20, 2021Updated 5 years ago
- A python package of robust and effective defogging/dehazing method☆15Dec 30, 2018Updated 7 years ago
- ☆34Jun 2, 2023Updated 3 years ago
- ☆29Nov 17, 2025Updated 9 months ago
- ☆17Jun 1, 2025Updated last year
- Cascade of CNNs for Robust Facial Landmarks Detection☆15Jan 29, 2021Updated 5 years ago
- ☆21Nov 30, 2019Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The implementation of FINER-MLLM, which is accepted by MM2024.☆18Oct 8, 2024Updated last year
- A curated list of Story Ending Generation models; DASFAA'22: Incorporating Commonsense Knowledge into Story Ending Generation via Heterog…☆15May 12, 2022Updated 4 years ago
- Rust standalone inference of Namo-500M series models. Extremly tiny, runing VLM on CPU.☆24Mar 12, 2025Updated last year
- Multimodal Variational Auto-encoder based Audio-Visual Segmentation [ICCV2023].☆20Sep 19, 2024Updated last year
- code repo for LoCoNet: Long-Short Context Network for Active Speaker Detection☆59May 1, 2023Updated 3 years ago
- AI agents weaving intelligence, execution, and automation into DeFi☆10Feb 28, 2025Updated last year
- ☆17Sep 27, 2020Updated 5 years ago
- Visual Speech Recognition for Multiple Languages☆480Aug 17, 2023Updated 3 years ago
- (ICLR 2025) Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation☆16Apr 29, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- We propose MMAD, a novel automated pipeline for precise AD generation. MMAD introduces ambient music alongside visual and linguistic, enh…☆17Dec 31, 2024Updated last year
- ☆24Sep 20, 2024Updated last year
- This is a fork of the original fairseq repository (version 0.12.2) with added classes for training mHuBERT-147.☆21Nov 19, 2024Updated last year
- Rank Centrality: Ranking from Pairwise Comparisons (Negahban et al 2016) implemented in Python☆17Oct 10, 2018Updated 7 years ago
- MultiOCR, an interface that connects multiple open-source OCR and various Cloud OCR.☆32Aug 19, 2023Updated 2 years ago
- ☆31Jul 17, 2026Updated last month
- Twitter bot for generating photo descriptions (alt text)☆23Jul 1, 2021Updated 5 years ago