[Interspeech 2026] Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness
☆21Jun 25, 2026Updated 3 weeks ago
Alternatives and similar repositories for UniTalk-ASD-code
Users that are interested in UniTalk-ASD-code are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [WACV 2026 Oral] LASER: Lip Landmark Assisted Speaker Detection for Robustness official implemntation☆30Feb 26, 2026Updated 4 months ago
- Official implementation of TalkNCE (ICASSP 2024).☆18Apr 30, 2025Updated last year
- code repo for LoCoNet: Long-Short Context Network for Active Speaker Detection☆57May 1, 2023Updated 3 years ago
- 🌸 A collection of Vietnamese women who are currently working in the field of Computer Science.☆16Jul 8, 2026Updated last week
- ☆19Jul 23, 2025Updated 11 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- The repository for Springer IJCV 2025 (LR-ASD: Lightweight and Robust Network for Active Speaker Detection)☆131Mar 23, 2025Updated last year
- python scripts for crawling original image from Google Images☆24May 5, 2022Updated 4 years ago
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Feb 11, 2026Updated 5 months ago
- Adaptive Flow-Matching for Target Speaker Extraction☆39Jul 13, 2026Updated last week
- INTERSPEECH2023: Target Active Speaker Detection with Audio-visual Cues☆61May 29, 2023Updated 3 years ago
- ☆35Aug 22, 2024Updated last year
- ☆20Nov 3, 2021Updated 4 years ago
- ☆27Jul 15, 2024Updated 2 years ago
- ☆17Apr 9, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆20Mar 20, 2026Updated 4 months ago
- repo for active speaker detection for media videos.☆31Nov 19, 2023Updated 2 years ago
- [CVPR'26, Findings] AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting☆15May 18, 2026Updated 2 months ago
- Collection of scripts from mHuBERT-147.☆35Nov 19, 2024Updated last year
- [NeurIPS 2025] AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding☆27Nov 3, 2025Updated 8 months ago
- pytorch implementation of PRNet, with weight transfered☆46Mar 6, 2019Updated 7 years ago
- Official implementation of paper "Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments"☆20Sep 25, 2025Updated 9 months ago
- ☆11Oct 2, 2024Updated last year
- Official Github repository for the CVPR 2022 paper "GIRAFFE HD: A High-Resolution 3D-aware Generative Model"☆70Nov 8, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official code for the paper "Scaling Multilingual Visual Speech Recognition"☆20Aug 15, 2025Updated 11 months ago
- Character-aware audio-only subtitling☆31Jun 15, 2025Updated last year
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- ADAPTING SELF-SUPERVISED MODELS TO MULTI-TALKER SPEECH RECOGNITION USING SPEAKER EMBEDDINGS☆33Mar 16, 2023Updated 3 years ago
- [INTERSPEECH 2022] This dataset is designed for multi-modal speaker diarization and lip-speech synchronization in the wild.☆65Jan 24, 2024Updated 2 years ago
- Look Who’s Talking: Active Speaker Detection in the Wild☆76Aug 24, 2023Updated 2 years ago
- Audio-Visual Speech Recognition☆26Jul 7, 2025Updated last year
- ☆18Jan 20, 2025Updated last year
- This repository contains the baseline system for CHiME-8 MMCSG challenge focusing on transcribing both sides of a conversation where one …☆41Mar 13, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for the paper: Graph Jigsaw Learning for Cartoon Face Recognition☆10Jul 1, 2022Updated 4 years ago
- ☆26Jun 10, 2026Updated last month
- 为视障人群生成电影,输入是电影剧本和mkv格式电影,输出为带有解说的电影☆12Jul 28, 2019Updated 6 years ago
- unofficial code of the paper "Efficient Geometry-aware 3D Generative Adversarial Networks"☆85Apr 9, 2022Updated 4 years ago
- ☆37Feb 7, 2026Updated 5 months ago
- ☆15Feb 18, 2024Updated 2 years ago
- Coloring lips and drawing glasses on faces in custom images or live webcam☆11Sep 10, 2019Updated 6 years ago