[WACV 2026 Oral] LASER: Lip Landmark Assisted Speaker Detection for Robustness official implemntation
☆30Feb 26, 2026Updated 6 months ago
Alternatives and similar repositories for LASER_ASD
Users that are interested in LASER_ASD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [Interspeech 2026] Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness☆24Jun 25, 2026Updated 2 months ago
- Official implementation of TalkNCE (ICASSP 2024).☆18Apr 30, 2025Updated last year
- The repository for Springer IJCV 2025 (LR-ASD: Lightweight and Robust Network for Active Speaker Detection)☆139Mar 23, 2025Updated last year
- ☆13Aug 24, 2023Updated 3 years ago
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Aug 4, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- INTERSPEECH2023: Target Active Speaker Detection with Audio-visual Cues☆61May 29, 2023Updated 3 years ago
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs☆17Jul 6, 2025Updated last year
- ☆19Jul 23, 2025Updated last year
- Animation of an SMPLX character in an augmented reality application☆19Aug 22, 2024Updated 2 years ago
- The project page repo for Neural Dubber.☆30Sep 20, 2023Updated 2 years ago
- ☆25Dec 19, 2024Updated last year
- ☆11Oct 2, 2024Updated last year
- Data manipulation and transformation for audio signal processing, powered by PyTorch☆10Sep 30, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code for Audio-Visual Target Speaker Extraction with Selective Auditory Attention (TASLP)☆35Feb 28, 2025Updated last year
- Dressed Human Reconstrcution from Single-view Real World Image☆25Mar 25, 2024Updated 2 years ago
- The repository for IEEE CVPR 2023 (A Light Weight Model for Active Speaker Detection)☆188Mar 23, 2025Updated last year
- PyTorch implementation of "Lip to Speech Synthesis with Visual Context Attentional GAN" (NeurIPS2021)☆25Mar 9, 2024Updated 2 years ago
- instagram reels in your terminal☆28Updated this week
- This repository contains the official implementation and pretrained weights for the paper "ReDimNet2: Scaling Speaker Verification via Ti…☆86Aug 28, 2026Updated last week
- This is the official implementation of work HiM2SAM in PRCV25.☆30Aug 30, 2025Updated last year
- Official implementation of A cappella: Audio-visual Singing VoiceSeparation, from BMVC21☆18May 14, 2022Updated 4 years ago
- ☆21Mar 4, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- SyncNet for Time Synchronization☆30Mar 13, 2023Updated 3 years ago
- ☆70Sep 13, 2022Updated 3 years ago
- Implementation of "Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition"☆18Aug 26, 2026Updated last week
- ☆12Dec 19, 2020Updated 5 years ago
- [INTERSPEECH 2022] This dataset is designed for multi-modal speaker diarization and lip-speech synchronization in the wild.☆67Jan 24, 2024Updated 2 years ago
- ☆13Jul 20, 2022Updated 4 years ago
- Coloring lips and drawing glasses on faces in custom images or live webcam☆11Sep 10, 2019Updated 6 years ago
- Audio-driven facial animation generator with BiLSTM used for transcribing the speech and web interface displaying the avatar and the anim…☆36Jul 14, 2022Updated 4 years ago
- Official repo of the paper “AL-GTD: Deep Active Learning for Gaze Target Detection” (ACMMM2024)☆12Jul 23, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official code for the paper "GestSync: Determining who is speaking without a talking head" published at BMVC 2023☆49Sep 1, 2024Updated 2 years ago
- Task-Focused Few-Shot Object Detection Benchmark☆14Jun 24, 2025Updated last year
- ☆25Jan 27, 2026Updated 7 months ago
- This branch of Asteroid contains code for the vocal harmony and chamber ensemble separation related papers.☆12Nov 7, 2024Updated last year
- ACM MM 2021: 'Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection'☆498Oct 23, 2023Updated 2 years ago
- Modified Python3 P2FA for Mandarin☆10Sep 21, 2020Updated 5 years ago
- Implementation of Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement☆18Oct 26, 2025Updated 10 months ago