[WACV 2026 Oral] LASER: Lip Landmark Assisted Speaker Detection for Robustness official implemntation
☆30Feb 26, 2026Updated 5 months ago
Alternatives and similar repositories for LASER_ASD
Users that are interested in LASER_ASD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [Interspeech 2026] Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness☆23Jun 25, 2026Updated last month
- Official implementation of TalkNCE (ICASSP 2024).☆18Apr 30, 2025Updated last year
- The repository for Springer IJCV 2025 (LR-ASD: Lightweight and Robust Network for Active Speaker Detection)☆137Mar 23, 2025Updated last year
- ☆13Aug 24, 2023Updated 2 years ago
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Aug 4, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- INTERSPEECH2023: Target Active Speaker Detection with Audio-visual Cues☆61May 29, 2023Updated 3 years ago
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs☆17Jul 6, 2025Updated last year
- [NeurIPS 2025] AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding☆27Nov 3, 2025Updated 9 months ago
- ☆19Jul 23, 2025Updated last year
- Animation of an SMPLX character in an augmented reality application☆19Aug 22, 2024Updated last year
- The project page repo for Neural Dubber.☆30Sep 20, 2023Updated 2 years ago
- ☆11Oct 2, 2024Updated last year
- Data manipulation and transformation for audio signal processing, powered by PyTorch☆10Sep 30, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Code for Audio-Visual Target Speaker Extraction with Selective Auditory Attention (TASLP)☆33Feb 28, 2025Updated last year
- ☆22Nov 24, 2022Updated 3 years ago
- 🍑 relsim: Relational Visual Similarity | pip install relsim 🌍 (CVPR 2026)☆87Jul 22, 2026Updated 3 weeks ago
- Video summarization using Vision Transformers☆14Feb 5, 2023Updated 3 years ago
- Dressed Human Reconstrcution from Single-view Real World Image☆25Mar 25, 2024Updated 2 years ago
- Self Reproduction Code of Paper "Reducing Transformer Key-Value Cache Size with Cross-Layer Attention (MIT CSAIL)☆17May 24, 2024Updated 2 years ago
- The repository for IEEE CVPR 2023 (A Light Weight Model for Active Speaker Detection)☆186Mar 23, 2025Updated last year
- PKU-Reid dataset☆14Jan 20, 2026Updated 6 months ago
- PyTorch implementation of "Lip to Speech Synthesis with Visual Context Attentional GAN" (NeurIPS2021)☆25Mar 9, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This repository contains the official implementation and pretrained weights for the paper "ReDimNet2: Scaling Speaker Verification via Ti…☆76Jul 30, 2026Updated 2 weeks ago
- ☆21Mar 4, 2024Updated 2 years ago
- SyncNet for Time Synchronization☆30Mar 13, 2023Updated 3 years ago
- ☆69Sep 13, 2022Updated 3 years ago
- Implementation of "Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition"☆16Aug 4, 2026Updated last week
- ☆12Dec 19, 2020Updated 5 years ago
- [INTERSPEECH 2022] This dataset is designed for multi-modal speaker diarization and lip-speech synchronization in the wild.☆67Jan 24, 2024Updated 2 years ago
- Audio-driven facial animation generator with BiLSTM used for transcribing the speech and web interface displaying the avatar and the anim…☆36Jul 14, 2022Updated 4 years ago
- Official repo of the paper “AL-GTD: Deep Active Learning for Gaze Target Detection” (ACMMM2024)☆12Jul 23, 2026Updated 3 weeks ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Task-Focused Few-Shot Object Detection Benchmark☆14Jun 24, 2025Updated last year
- ☆11Nov 5, 2021Updated 4 years ago
- This branch of Asteroid contains code for the vocal harmony and chamber ensemble separation related papers.☆12Nov 7, 2024Updated last year
- ACM MM 2021: 'Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection'☆495Oct 23, 2023Updated 2 years ago
- Modified Python3 P2FA for Mandarin☆10Sep 21, 2020Updated 5 years ago
- 重庆理工大学本科毕业设计 题目:基于深度学习的行人重识别☆16Apr 10, 2023Updated 3 years ago
- Implementation of Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement☆18Oct 26, 2025Updated 9 months ago