A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation Pattern
☆31Mar 24, 2026Updated 4 months ago
Alternatives and similar repositories for Spoken-Dialogue-Model-Survey
Users that are interested in Spoken-Dialogue-Model-Survey are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM-Codec: Neural Audio Codec Meets Language Model Objectives☆23May 3, 2026Updated 3 months ago
- AudioCodec-Hub is a Python library for encoding and decoding audio data, supporting various neural audio codec models☆25Sep 26, 2023Updated 2 years ago
- Code and model for ICASSP 2025 Paper "Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data"☆127Jul 15, 2025Updated last year
- Super Flappy Bird in p5.js☆10Mar 8, 2021Updated 5 years ago
- Unsupervised spoken sentence embeddings☆14Dec 14, 2022Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- The official repository of Dynamic-SUPERB.☆200Jun 24, 2025Updated last year
- REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR☆15Dec 11, 2024Updated last year
- Understanding and Tackling Hallucinations in Large Audio-Language Models | ICASSP 2025, Interspeech 2024☆34Mar 14, 2025Updated last year
- ☆15Sep 9, 2021Updated 4 years ago
- Yeast, a lite and light beamer theme☆18Dec 6, 2020Updated 5 years ago
- ☆10Sep 19, 2022Updated 3 years ago
- A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models (ICASSP 2024)☆58Apr 17, 2024Updated 2 years ago
- A PyTorch implementation of the universal neural vocoder☆68Nov 6, 2020Updated 5 years ago
- Lightweight streaming Voice Activity Detection (VAD) tool with ONNX runtime☆23Mar 18, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Some PKGBUILDs☆12Aug 5, 2025Updated last year
- Official repository for the paper Multimodal Transformer Distillation for Audio-Visual Synchronization (ICASSP 2024).☆29Apr 3, 2024Updated 2 years ago
- Audio Codec Speech processing Universal PERformance Benchmark☆309Jul 4, 2026Updated last month
- Code for DeSTA2.5-Audio, general-purpose LALM☆141Feb 4, 2026Updated 6 months ago
- COG-MHEAR Audio-Visual Speech Enhancement Challenge☆48Feb 17, 2026Updated 5 months ago
- Collection of works for evaluating (and analyzing) large audio-language models (LALMs)☆41Aug 11, 2025Updated 11 months ago
- ASR text preprocessing utility☆21Aug 5, 2024Updated 2 years ago
- Toward Multi Modality Language Model - implementation of GPT-4o/Project Astra☆16Dec 10, 2024Updated last year
- ☆23Aug 21, 2020Updated 5 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis☆69Jun 12, 2026Updated last month
- [ACL 2026 Main] MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows☆147Sep 2, 2025Updated 11 months ago
- Official Repository of UltraVoice☆63Oct 28, 2025Updated 9 months ago
- fd-sds☆21Apr 8, 2026Updated 4 months ago
- [ICASSP 2026] Official code for "Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration"☆17Apr 16, 2026Updated 3 months ago
- A curated list of full-duplex spoken dialogue models & benchmarks☆172Updated this week
- Official implementation: "AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation"☆20Oct 9, 2025Updated 10 months ago
- Official Implementation and Dataset of paper - DFADD: The Diffusion and Flow-matching based Audio Deepfake Dataset☆16Apr 7, 2025Updated last year
- Official PyTorch implementation of "RVAE-EM: Generative speech dereverberation based on recurrent variational auto-encoder and convolutiv…☆51Mar 6, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 《SpeechPrompt v2: Prompt Tuning for Speech Classification Tasks》Speech processing with prompting paradigm☆81Oct 19, 2023Updated 2 years ago
- The accompanying code for "Exploring the limits of decoder-only models trained on public speech recognition corpora" (Ankit Gupta, George…☆21Oct 11, 2024Updated last year
- NVV-SuperBench: Beyond Words, Beyond Quality—Benchmarking Nonverbal Vocalizations in Speech Generation (Interspeech 2026 long paper)☆18Jun 21, 2026Updated last month
- A Benchmark for Evaluating Turn-Taking and Overlap Handling in Full-Duplex Spoken Dialogue Models☆256May 20, 2026Updated 2 months ago
- Unofficial fairseq-free PyTorch implementation of UTMOS (v1, 2022), matching the original system.☆35Jun 6, 2026Updated 2 months ago
- ☆29Jul 25, 2026Updated 2 weeks ago
- This repository contains the code for the paper "voc2vec: A Foundation Model for Non-Verbal Vocalization", accepted at ICASSP 2025.☆58Apr 14, 2025Updated last year