Official repository for the paper Multimodal Transformer Distillation for Audio-Visual Synchronization (ICASSP 2024).
☆29Apr 3, 2024Updated 2 years ago
Alternatives and similar repositories for MTDVocaLiST
Users that are interested in MTDVocaLiST are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for the paper VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices☆73Apr 7, 2024Updated 2 years ago
- Official Implementation and Dataset of paper - DFADD: The Diffusion and Flow-matching based Audio Deepfake Dataset☆16Apr 7, 2025Updated last year
- A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation …☆31Mar 24, 2026Updated 4 months ago
- This repository collects papers related to Speech Tokenizer.☆18Oct 16, 2024Updated last year
- Unsupervised spoken sentence embeddings☆14Dec 14, 2022Updated 3 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Official repository for the paper Singing Voice Graph Modeling for SingFake Detection (Interspeech 2024).☆24Sep 19, 2025Updated 10 months ago
- EMO-SUPERB: a reproducible speech emotion recognition benchmark with leakage-free splits for 6 datasets and 15 speech SSL models (IEEE SL…☆51Updated this week
- awesome-audio-visual-robustness☆11Jan 27, 2024Updated 2 years ago
- 《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》☆77Jun 9, 2023Updated 3 years ago
- [EMNLP 2024] ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers☆126Mar 20, 2025Updated last year
- The open source code of ALMTokenizer2: Towards Low bit-rate and Semantic-rich Audio Tokenizer with Flow-based Scalar Diffusion Transforme…☆45Sep 5, 2025Updated 10 months ago
- A fully and partially fake speech dataset for evaluation☆15Nov 11, 2025Updated 8 months ago
- A casual and simple ChatGPT Python script that can run using terminal (as long as you have an API). Support Azure API.☆20May 3, 2025Updated last year
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.☆13Oct 11, 2022Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [Interspeech 2025] DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec☆72Mar 11, 2026Updated 4 months ago
- ☆12Mar 28, 2024Updated 2 years ago
- Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis☆40Sep 14, 2023Updated 2 years ago
- Config files for my GitHub profile.☆11Updated this week
- LLM-Codec: Neural Audio Codec Meets Language Model Objectives☆23May 3, 2026Updated 2 months ago
- This repository presents FSD dataset for song deepfake detection.☆24Aug 18, 2025Updated 11 months ago
- The demo page for ALMTokenizer☆59Apr 14, 2025Updated last year
- R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning☆82Jan 3, 2024Updated 2 years ago
- NSA Cybersecurity. Formerly known as NSA Information Assurance and the Information Assurance Directorate☆10Jul 7, 2022Updated 4 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- This is the official supplementary document for the GCT data and its prediction task.☆10Feb 19, 2024Updated 2 years ago
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- Inference code for Interspeech 2025 paper, "LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec"☆36Oct 23, 2025Updated 9 months ago
- AudioCodec-Hub is a Python library for encoding and decoding audio data, supporting various neural audio codec models☆25Sep 26, 2023Updated 2 years ago
- ☆35Sep 6, 2025Updated 10 months ago
- ☆10Feb 17, 2023Updated 3 years ago
- Modeling Harmonic Complexity using two models of Conditional Variational Autoencoders - MSc. Thesis☆10May 16, 2023Updated 3 years ago
- Official Repository for "SingFake: Singing Voice Deepfake Detection"☆64Feb 26, 2024Updated 2 years ago
- A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models (ICASSP 2024)☆58Apr 17, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- 🦅🔗 Building FlyteGPT on Flyte with LangChain☆30Jan 23, 2024Updated 2 years ago
- Audio Research in US. US-based professors who work on audio (music, speech, acoustics). For students who would like to apply for RA, PhD,…☆27Feb 27, 2026Updated 5 months ago
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.☆125Jun 4, 2025Updated last year
- IEEE VCIP 2021: AnomalyHop: An SSL-based Image Anomaly Localization Method☆14Sep 18, 2021Updated 4 years ago
- Github repository for ACL 2025 paper: VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models☆24Jun 16, 2025Updated last year
- About Us☆19Mar 30, 2024Updated 2 years ago
- ☆10Aug 23, 2021Updated 4 years ago