Official repository for the paper Multimodal Transformer Distillation for Audio-Visual Synchronization (ICASSP 2024).
☆29Apr 3, 2024Updated 2 years ago
Alternatives and similar repositories for MTDVocaLiST
Users that are interested in MTDVocaLiST are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for the paper VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices☆73Apr 7, 2024Updated 2 years ago
- Official Implementation and Dataset of paper - DFADD: The Diffusion and Flow-matching based Audio Deepfake Dataset☆17Apr 7, 2025Updated last year
- A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation …☆32Mar 24, 2026Updated 6 months ago
- This repository collects papers related to Speech Tokenizer.☆18Oct 16, 2024Updated last year
- Unsupervised spoken sentence embeddings☆14Dec 14, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Official repository for the paper Singing Voice Graph Modeling for SingFake Detection (Interspeech 2024).☆24Sep 19, 2025Updated last year
- EMO-SUPERB: a reproducible speech emotion recognition benchmark with leakage-free splits for 6 datasets and 15 speech SSL models (IEEE SL…☆52Jul 30, 2026Updated last month
- awesome-audio-visual-robustness☆12Jan 27, 2024Updated 2 years ago
- 《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》☆77Jun 9, 2023Updated 3 years ago
- [EMNLP 2024] ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers☆127Mar 20, 2025Updated last year
- A fully and partially fake speech dataset for evaluation☆15Nov 11, 2025Updated 10 months ago
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.☆14Oct 11, 2022Updated 3 years ago
- A casual and simple ChatGPT Python script that can run using terminal (as long as you have an API). Support Azure API.☆20May 3, 2025Updated last year
- [Interspeech 2025] DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec☆73Sep 9, 2026Updated 2 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis☆41Sep 14, 2023Updated 3 years ago
- Config files for my GitHub profile.☆12Jul 25, 2026Updated 2 months ago
- LLM-Codec: Neural Audio Codec Meets Language Model Objectives☆23May 3, 2026Updated 4 months ago
- A Transformer approach for polyphonic Audio-to-Score (A2S) transcription (ICASSP 2024)☆15Jun 21, 2025Updated last year
- ATTENTION AGGREGATION NETWORK FOR AUDIO-VISUAL EMOTION RECOGNITION☆14Sep 25, 2023Updated 3 years ago
- This repository presents FSD dataset for song deepfake detection.☆24Aug 18, 2025Updated last year
- The demo page for ALMTokenizer☆59Apr 14, 2025Updated last year
- This repository contains the code for the paper "voc2vec: A Foundation Model for Non-Verbal Vocalization", accepted at ICASSP 2025.☆59Apr 14, 2025Updated last year
- R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning☆82Jan 3, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- NSA Cybersecurity. Formerly known as NSA Information Assurance and the Information Assurance Directorate☆10Jul 7, 2022Updated 4 years ago
- This is the official supplementary document for the GCT data and its prediction task.☆10Feb 19, 2024Updated 2 years ago
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- Inference code for Interspeech 2025 paper, "LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec"☆37Oct 23, 2025Updated 11 months ago
- AudioCodec-Hub is a Python library for encoding and decoding audio data, supporting various neural audio codec models☆25Sep 26, 2023Updated 3 years ago
- ☆35Sep 6, 2025Updated last year
- ☆10Feb 17, 2023Updated 3 years ago
- Modeling Harmonic Complexity using two models of Conditional Variational Autoencoders - MSc. Thesis☆10May 16, 2023Updated 3 years ago
- Official Repository for "SingFake: Singing Voice Deepfake Detection"☆65Feb 26, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 🦅🔗 Building FlyteGPT on Flyte with LangChain☆30Jan 23, 2024Updated 2 years ago
- A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models (ICASSP 2024)☆58Apr 17, 2024Updated 2 years ago
- The official implementation of ImageBind-LLM and Whisper-LLM from the paper "Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Compre…☆20Oct 30, 2023Updated 2 years ago
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.☆129Jun 4, 2025Updated last year
- IEEE VCIP 2021: AnomalyHop: An SSL-based Image Anomaly Localization Method☆14Sep 18, 2021Updated 5 years ago
- Github repository for ACL 2025 paper: VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models☆25Jun 16, 2025Updated last year
- About Us☆19Mar 30, 2024Updated 2 years ago