[WACV 2026 Oral] LASER: Lip Landmark Assisted Speaker Detection for Robustness official implemntation
☆30Feb 26, 2026Updated 6 months ago
Alternatives and similar repositories for LASER_ASD
Users that are interested in LASER_ASD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of TalkNCE (ICASSP 2024).☆19Apr 30, 2025Updated last year
- code repo for LoCoNet: Long-Short Context Network for Active Speaker Detection☆61May 1, 2023Updated 3 years ago
- The repository for Springer IJCV 2025 (LR-ASD: Lightweight and Robust Network for Active Speaker Detection)☆144Mar 23, 2025Updated last year
- ☆13Aug 24, 2023Updated 3 years ago
- Official Code for the CVPR 2026 Paper "MATCH: Feed-forward Gaussian Registration for Head Avatar Creation and Editing"☆22Jun 6, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- INTERSPEECH2023: Target Active Speaker Detection with Audio-visual Cues☆63May 29, 2023Updated 3 years ago
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs☆17Jul 6, 2025Updated last year
- Using this repository, you can identify people in the video by data into a SQLite database, and re-identifying them whenever they appear …☆23Jun 22, 2022Updated 4 years ago
- ☆19Jul 23, 2025Updated last year
- [NeurIPS 2025] AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding☆27Nov 3, 2025Updated 10 months ago
- Animation of an SMPLX character in an augmented reality application☆19Aug 22, 2024Updated 2 years ago
- The project page repo for Neural Dubber.☆30Sep 20, 2023Updated 3 years ago
- ☆25Dec 19, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆11Oct 2, 2024Updated last year
- Code for Audio-Visual Target Speaker Extraction with Selective Auditory Attention (TASLP)☆35Feb 28, 2025Updated last year
- ☆22Nov 24, 2022Updated 3 years ago
- Identity-preserving Distillation Sampling for Image Translation with Latent Diffusion Models☆17May 15, 2025Updated last year
- Dressed Human Reconstrcution from Single-view Real World Image☆25Mar 25, 2024Updated 2 years ago
- The repository for IEEE CVPR 2023 (A Light Weight Model for Active Speaker Detection)☆191Mar 23, 2025Updated last year
- PyTorch implementation of "Lip to Speech Synthesis with Visual Context Attentional GAN" (NeurIPS2021)☆25Mar 9, 2024Updated 2 years ago
- Official implementation of A cappella: Audio-visual Singing VoiceSeparation, from BMVC21☆18May 14, 2022Updated 4 years ago
- ☆21Mar 4, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Visual Speech Recognition for Multiple Languages☆483Aug 17, 2023Updated 3 years ago
- Speaker-aware CTC (SACTC) for multi-talker overlapped speech recognition.☆22May 26, 2025Updated last year
- SyncNet for Time Synchronization☆30Mar 13, 2023Updated 3 years ago
- ☆70Sep 13, 2022Updated 4 years ago
- The inference code of RVC-Boss/GPT-SoVITS that can be developer-friendly.☆16Sep 29, 2024Updated last year
- ☆16Jan 13, 2024Updated 2 years ago
- [INTERSPEECH 2022] This dataset is designed for multi-modal speaker diarization and lip-speech synchronization in the wild.☆69Jan 24, 2024Updated 2 years ago
- ☆13Jul 20, 2022Updated 4 years ago
- Coloring lips and drawing glasses on faces in custom images or live webcam☆11Sep 10, 2019Updated 7 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Audio-driven facial animation generator with BiLSTM used for transcribing the speech and web interface displaying the avatar and the anim…☆36Jul 14, 2022Updated 4 years ago
- Official repo of the paper “AL-GTD: Deep Active Learning for Gaze Target Detection” (ACMMM2024)☆12Jul 23, 2026Updated 2 months ago
- Official code for the paper "GestSync: Determining who is speaking without a talking head" published at BMVC 2023☆49Sep 1, 2024Updated 2 years ago
- Task-Focused Few-Shot Object Detection Benchmark☆14Jun 24, 2025Updated last year
- PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification☆52Jun 11, 2025Updated last year
- This branch of Asteroid contains code for the vocal harmony and chamber ensemble separation related papers.☆12Nov 7, 2024Updated last year
- ACM MM 2021: 'Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection'☆505Oct 23, 2023Updated 2 years ago