Disentangled Speech Embeddings using Cross-Modal Self-Supervision
☆167Apr 12, 2020Updated 6 years ago
Alternatives and similar repositories for syncnet_trainer
Users that are interested in syncnet_trainer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Out of time: automated lip sync in the wild☆894Apr 17, 2026Updated 3 months ago
- Augmentation adversarial training for self-supervised speaker recognition☆77Aug 15, 2021Updated 4 years ago
- In defence of metric learning for speaker recognition☆1,170Apr 22, 2026Updated 2 months ago
- Development Toolkit for the VoxCeleb Speaker Recognition Challenge 2020☆43Jul 17, 2020Updated 6 years ago
- the dataset and code for "Flow-guided One-shot Talking Face Generation with a High-resolution Audio-visual Dataset"☆429May 12, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Utterance-level Aggregation For Speaker Recognition In The Wild☆371Mar 24, 2023Updated 3 years ago
- Implementation for ECCV20 paper "Self-Supervised Learning of audio-visual objects from video"☆114Nov 16, 2020Updated 5 years ago
- Optimized Syncnet and Chinese enhanced version, EN and CN checkpoints released☆11Nov 8, 2021Updated 4 years ago
- Official repository for the paper VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices☆73Apr 7, 2024Updated 2 years ago
- Code for Audio-Visual Target Speaker Extraction with Selective Auditory Attention (TASLP)☆32Feb 28, 2025Updated last year
- A PyTorch implementation of the Deep Audio-Visual Speech Recognition paper.☆244Feb 15, 2024Updated 2 years ago
- ☆42Nov 22, 2024Updated last year
- [InterSpeech 2020] "AutoSpeech: Neural Architecture Search for Speaker Recognition" by Shaojin Ding*, Tianlong Chen*, Xinyu Gong, Weiwei …☆206Dec 8, 2022Updated 3 years ago
- Audio-visual diarization pipeline used for creating VoxConverse dataset☆22Jun 6, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A light weight neural speaker embeddings extraction based on Kaldi and PyTorch.☆136Jan 27, 2020Updated 6 years ago
- PyTorch implementation of RPNSD☆60Jun 17, 2024Updated 2 years ago
- ☆21Apr 6, 2021Updated 5 years ago
- Unsupervised Speech Decomposition via Triple Information Bottleneck☆14Apr 29, 2020Updated 6 years ago
- ☆105Jul 5, 2023Updated 3 years ago
- Python implementation of the paper " Dynamic Temporal Alignment of Speech to Lips"☆32May 16, 2019Updated 7 years ago
- SyncNet for Time Synchronization☆30Mar 13, 2023Updated 3 years ago
- Code for "Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose" (Arxiv 2020) and "Predicting Personalize…☆772Dec 15, 2023Updated 2 years ago
- Demo for 2022 Interspeech☆29Jun 14, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆838Nov 19, 2025Updated 8 months ago
- ☆18Nov 22, 2024Updated last year
- Code and instruction on replicating the experiments done in paper: Unified Hypersphere Embedding for Speaker Recognition☆32Jul 14, 2019Updated 7 years ago
- A self-supervised learning framework for audio-visual speech☆992Dec 7, 2023Updated 2 years ago
- Audio-Visual Speech Separation with Cross-Modal Consistency☆250Jul 25, 2023Updated 2 years ago
- Tensorflow implementation of x-vector topology on top of Kaldi recipe☆118Nov 5, 2019Updated 6 years ago
- ☆65Jun 28, 2023Updated 3 years ago
- Real-time melgan based on cpu !!!☆13Dec 3, 2019Updated 6 years ago
- PyTorch implementation of "StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator"☆215Aug 8, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official github repo for paper "What comprises a good talking-head video generation?: A Survey and Benchmark"☆91Dec 8, 2022Updated 3 years ago
- MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation [ECCV2020]☆305Jul 7, 2024Updated 2 years ago
- ☆17Aug 27, 2025Updated 10 months ago
- Pytorch implementation of Generalized End-to-End Loss for speaker verification☆88Apr 23, 2019Updated 7 years ago
- Deep speaker embeddings in PyTorch, including x-vectors. Code used in this work: https://arxiv.org/abs/2007.16196☆321Nov 11, 2020Updated 5 years ago
- ☆429Nov 1, 2023Updated 2 years ago
- ☆15Apr 6, 2023Updated 3 years ago