Official repository for the paper Multimodal Transformer Distillation for Audio-Visual Synchronization (ICASSP 2024).
☆29Apr 3, 2024Updated 2 years ago
Alternatives and similar repositories for MTDVocaLiST
Users that are interested in MTDVocaLiST are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for the paper VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices☆73Apr 7, 2024Updated 2 years ago
- Official Implementation and Dataset of paper - DFADD: The Diffusion and Flow-matching based Audio Deepfake Dataset☆16Apr 7, 2025Updated last year
- A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation …☆32Mar 24, 2026Updated 4 months ago
- This repository collects papers related to Speech Tokenizer.☆18Oct 16, 2024Updated last year
- Unsupervised spoken sentence embeddings☆14Dec 14, 2022Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Official repository for the paper Singing Voice Graph Modeling for SingFake Detection (Interspeech 2024).☆24Sep 19, 2025Updated 10 months ago
- EMO-SUPERB: a reproducible speech emotion recognition benchmark with leakage-free splits for 6 datasets and 15 speech SSL models (IEEE SL…☆52Jul 30, 2026Updated 2 weeks ago
- awesome-audio-visual-robustness☆11Jan 27, 2024Updated 2 years ago
- 《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》☆77Jun 9, 2023Updated 3 years ago
- [EMNLP 2024] ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers☆127Mar 20, 2025Updated last year
- The open source code of ALMTokenizer2: Towards Low bit-rate and Semantic-rich Audio Tokenizer with Flow-based Scalar Diffusion Transforme…☆45Sep 5, 2025Updated 11 months ago
- A fully and partially fake speech dataset for evaluation☆15Nov 11, 2025Updated 9 months ago
- A casual and simple ChatGPT Python script that can run using terminal (as long as you have an API). Support Azure API.☆20May 3, 2025Updated last year
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.☆13Oct 11, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [Interspeech 2025] DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec☆72Mar 11, 2026Updated 5 months ago
- ☆12Mar 28, 2024Updated 2 years ago
- Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis☆40Sep 14, 2023Updated 2 years ago
- LLM-Codec: Neural Audio Codec Meets Language Model Objectives☆23May 3, 2026Updated 3 months ago
- A Transformer approach for polyphonic Audio-to-Score (A2S) transcription (ICASSP 2024)☆15Jun 21, 2025Updated last year
- ATTENTION AGGREGATION NETWORK FOR AUDIO-VISUAL EMOTION RECOGNITION☆14Sep 25, 2023Updated 2 years ago
- This repository presents FSD dataset for song deepfake detection.☆24Aug 18, 2025Updated 11 months ago
- The demo page for ALMTokenizer☆59Apr 14, 2025Updated last year
- This repository contains the code for the paper "voc2vec: A Foundation Model for Non-Verbal Vocalization", accepted at ICASSP 2025.☆59Apr 14, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning☆82Jan 3, 2024Updated 2 years ago
- NSA Cybersecurity. Formerly known as NSA Information Assurance and the Information Assurance Directorate☆10Jul 7, 2022Updated 4 years ago
- This is the official supplementary document for the GCT data and its prediction task.☆10Feb 19, 2024Updated 2 years ago
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- Inference code for Interspeech 2025 paper, "LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec"☆36Oct 23, 2025Updated 9 months ago
- AudioCodec-Hub is a Python library for encoding and decoding audio data, supporting various neural audio codec models☆25Sep 26, 2023Updated 2 years ago
- ☆35Sep 6, 2025Updated 11 months ago
- ☆10Feb 17, 2023Updated 3 years ago
- Modeling Harmonic Complexity using two models of Conditional Variational Autoencoders - MSc. Thesis☆10May 16, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official Repository for "SingFake: Singing Voice Deepfake Detection"☆65Feb 26, 2024Updated 2 years ago
- A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models (ICASSP 2024)☆58Apr 17, 2024Updated 2 years ago
- 🦅🔗 Building FlyteGPT on Flyte with LangChain☆30Jan 23, 2024Updated 2 years ago
- The official implementation of ImageBind-LLM and Whisper-LLM from the paper "Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Compre…☆20Oct 30, 2023Updated 2 years ago
- Audio Research in US. US-based professors who work on audio (music, speech, acoustics). For students who would like to apply for RA, PhD,…☆29Feb 27, 2026Updated 5 months ago
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.☆126Jun 4, 2025Updated last year
- Github repository for ACL 2025 paper: VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models☆24Jun 16, 2025Updated last year