☆27Jun 27, 2023Updated 3 years ago
Alternatives and similar repositories for cncvs_data_collector
Users that are interested in cncvs_data_collector are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The MAVD represents Mandarin Audio-Visual dataset with Depth information. MAVD has a rich variety of modal data, including audio, RGB ima…☆20Apr 22, 2024Updated 2 years ago
- Baseline system for CNVSRC2023 (Chinese Continuous Visual Speech Recognition Challenge 2023)☆23Apr 27, 2024Updated 2 years ago
- Official CNVSRC2025 Competition Baseline☆17Jun 27, 2025Updated last year
- wav2lip in a Vector Quantized (VQ) space☆27Jun 20, 2023Updated 3 years ago
- FaceFormer Emo: Speech-Driven 3D Facial Animation with Emotion Embedding☆27Jul 15, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- An unofficial (PyTorch) implementation for the paper Deep Lip Reading: A comparison of models and an online application.☆10May 13, 2020Updated 6 years ago
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into one☆26Aug 5, 2024Updated 2 years ago
- wav2lip-api☆11Mar 16, 2023Updated 3 years ago
- Visual Speech Recongnition☆22Dec 24, 2024Updated last year
- Speech-Driven Expression Blendshape Based on Single-Layer Self-attention Network (AIWIN 2022)☆77Oct 21, 2022Updated 3 years ago
- optimized wav2lip☆18Jan 6, 2024Updated 2 years ago
- Identifying "who speak when" using visual speech input and pretrained lip-sync expert☆18Jul 1, 2023Updated 3 years ago
- ☆10Feb 17, 2023Updated 3 years ago
- ☆21Mar 4, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Implementation of "Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition"☆18Aug 26, 2026Updated last week
- ☆21Dec 9, 2023Updated 2 years ago
- Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language (AAAI 2025)☆24Jun 29, 2026Updated 2 months ago
- ☆56Dec 20, 2023Updated 2 years ago
- ☆104Nov 26, 2025Updated 9 months ago
- ☆10Nov 19, 2023Updated 2 years ago
- [IJCAI2022] Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast☆22Oct 25, 2023Updated 2 years ago
- Tsinghua University SPMI Lab array processing toolkit☆18Nov 23, 2016Updated 9 years ago
- [BMVC'24] G3FA: Geometry-guided GAN for Face Animation☆20Mar 14, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [CVPR 2026 Findings] TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis☆200Jun 8, 2026Updated 2 months ago
- Ego4DSounds: A diverse egocentric dataset with high action-audio correspondence☆21Jun 14, 2024Updated 2 years ago
- An implementation of http://openaccess.thecvf.com/content_CVPRW_2019/papers/Sight%20and%20Sound/Konstantinos_Vougioukas_End-to-End_Speech…☆18Mar 19, 2020Updated 6 years ago
- Prompting Large Language Models with Audio for General-Purpose Speech Summarization☆20May 14, 2025Updated last year
- SyncNet for Time Synchronization☆30Mar 13, 2023Updated 3 years ago
- ☆24Oct 8, 2021Updated 4 years ago
- Code for "Self-Lifting: A Novel Framework For Unsupervised Voice-Face Association Learning,ICMR,2022"☆15Oct 25, 2024Updated last year
- ☆15Oct 28, 2019Updated 6 years ago
- CVPR 2022: Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?☆130Dec 11, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Pushing the limits of acoustic motion tracking☆15Jul 31, 2020Updated 6 years ago
- MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation [ECCV2020]☆304Jul 7, 2024Updated 2 years ago
- 这是一个在wav2lip,使用wav2lip、gfpgan、yolov5等模型用RT加速的超快推理!经测试在2070显卡上可达到0.03秒每帧实现实时推理。☆32Sep 23, 2025Updated 11 months ago
- ☆16Apr 27, 2025Updated last year
- PaintsTorch: Automatic Lineart Colorization☆10Jun 21, 2019Updated 7 years ago
- Official implementation of SBNet as described in "Single-branch Network for Multimodal Training".☆13Aug 28, 2023Updated 3 years ago
- Auto-AVSR: Lip-Reading Sentences Project☆433Jan 8, 2025Updated last year