[WACV 2026 Oral] LASER: Lip Landmark Assisted Speaker Detection for Robustness official implemntation
☆30Feb 26, 2026Updated 5 months ago
Alternatives and similar repositories for LASER_ASD
Users that are interested in LASER_ASD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of TalkNCE (ICASSP 2024).☆18Apr 30, 2025Updated last year
- The repository for Springer IJCV 2025 (LR-ASD: Lightweight and Robust Network for Active Speaker Detection)☆132Mar 23, 2025Updated last year
- ☆13Aug 24, 2023Updated 2 years ago
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Feb 11, 2026Updated 5 months ago
- Official Code for the CVPR 2026 Paper "MATCH: Feed-forward Gaussian Registration for Head Avatar Creation and Editing"☆16Jun 6, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- INTERSPEECH2023: Target Active Speaker Detection with Audio-visual Cues☆61May 29, 2023Updated 3 years ago
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs☆17Jul 6, 2025Updated last year
- [NeurIPS 2025] AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding☆27Nov 3, 2025Updated 8 months ago
- ☆19Jul 23, 2025Updated last year
- Animation of an SMPLX character in an augmented reality application☆19Aug 22, 2024Updated last year
- The project page repo for Neural Dubber.☆30Sep 20, 2023Updated 2 years ago
- ☆25Dec 19, 2024Updated last year
- ☆11Oct 2, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Data manipulation and transformation for audio signal processing, powered by PyTorch☆10Sep 30, 2024Updated last year
- Code for Audio-Visual Target Speaker Extraction with Selective Auditory Attention (TASLP)☆32Feb 28, 2025Updated last year
- Dressed Human Reconstrcution from Single-view Real World Image☆25Mar 25, 2024Updated 2 years ago
- The repository for IEEE CVPR 2023 (A Light Weight Model for Active Speaker Detection)☆181Mar 23, 2025Updated last year
- PyTorch implementation of "Lip to Speech Synthesis with Visual Context Attentional GAN" (NeurIPS2021)☆25Mar 9, 2024Updated 2 years ago
- instagram reels in your terminal☆25Updated this week
- This repository contains the official implementation and pretrained weights for the paper "ReDimNet2: Scaling Speaker Verification via Ti…☆65Jul 9, 2026Updated 2 weeks ago
- This is the official implementation of work HiM2SAM in PRCV25.☆29Aug 30, 2025Updated 10 months ago
- Official implementation of A cappella: Audio-visual Singing VoiceSeparation, from BMVC21☆18May 14, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆21Mar 4, 2024Updated 2 years ago
- Visual Speech Recognition for Multiple Languages☆478Aug 17, 2023Updated 2 years ago
- Speaker-aware CTC (SACTC) for multi-talker overlapped speech recognition.☆22May 26, 2025Updated last year
- SyncNet for Time Synchronization☆30Mar 13, 2023Updated 3 years ago
- ☆69Sep 13, 2022Updated 3 years ago
- The inference code of RVC-Boss/GPT-SoVITS that can be developer-friendly.☆16Sep 29, 2024Updated last year
- Implementation of "Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition"☆16Jun 9, 2026Updated last month
- ☆11Dec 19, 2020Updated 5 years ago
- ☆16Jan 13, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Jul 20, 2022Updated 4 years ago
- FBX file loader for python (only supports geometry currently)☆17Aug 5, 2024Updated last year
- Audio-driven facial animation generator with BiLSTM used for transcribing the speech and web interface displaying the avatar and the anim…☆36Jul 14, 2022Updated 4 years ago
- Official repo of the paper “AL-GTD: Deep Active Learning for Gaze Target Detection” (ACMMM2024)☆12Updated this week
- Official code for the paper "GestSync: Determining who is speaking without a talking head" published at BMVC 2023☆48Sep 1, 2024Updated last year
- Task-Focused Few-Shot Object Detection Benchmark☆14Jun 24, 2025Updated last year
- ☆11Nov 5, 2021Updated 4 years ago