☆141Jan 24, 2026Updated 6 months ago
Alternatives and similar repositories for Speech-and-audio-papers-Top-Conference
Users that are interested in Speech-and-audio-papers-Top-Conference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PyTorch implementation of USR 2.0 (ICLR 2026)☆15Apr 3, 2026Updated 3 months ago
- Audio Codec Speech processing Universal PERformance Benchmark☆308Jul 4, 2026Updated 3 weeks ago
- 🦇 Encoder of BAT (Learning to Reason about Spatial Sounds with Large Language Models)☆87Feb 13, 2025Updated last year
- Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with syntheti…☆137Mar 3, 2026Updated 4 months ago
- Awesome speech/audio LLMs, representation learning, and codec models☆1,240Jul 10, 2026Updated 2 weeks ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- (ICLR 2025) Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation☆16Apr 29, 2025Updated last year
- Update ASR paper everyday☆513May 16, 2026Updated 2 months ago
- Understanding and Tackling Hallucinations in Large Audio-Language Models | ICASSP 2025, Interspeech 2024☆34Mar 14, 2025Updated last year
- Awesome Neural Codec Models, Text-to-Speech Synthesizers & Speech Language Models☆246Jul 9, 2026Updated 2 weeks ago
- ☆47Apr 2, 2025Updated last year
- [ACL 2025 Main] UniCodec: a unified audio codec with a single codebook to support multi-domain audio data, including speech, music, and s…☆157May 30, 2025Updated last year
- The implementation for "Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions"☆51Apr 7, 2025Updated last year
- VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling☆100Nov 9, 2024Updated last year
- Unofficial pytorch reproduction for the paper "Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction" (…☆60Apr 4, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)☆662Updated this week
- The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio gener…☆948Updated this week
- Evaluation Protocol for Large-Scale Zero-Shot TTS Literature☆97Mar 12, 2025Updated last year
- [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix☆214Feb 25, 2026Updated 4 months ago
- [ACL 2026 Main] MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows☆142Sep 2, 2025Updated 10 months ago
- Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'☆163Mar 26, 2026Updated 3 months ago
- LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement☆105Apr 1, 2025Updated last year
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling☆175Feb 28, 2025Updated last year
- [ICCV'25] Official PyTorch Implementation of "VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models"☆17Dec 8, 2025Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2024] AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation☆48Sep 6, 2024Updated last year
- [NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”☆38Jul 24, 2025Updated last year
- This is an evolving repo for the paper "Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey".☆274Updated this week
- The official repository of SpeechCraft dataset, a large-scale expressive bilingual speech dataset with natural language descriptions.☆197Feb 28, 2026Updated 4 months ago
- Curated list for papers, codes and resources related to Text-to-Audio (TTA) Generation☆74Updated this week
- Official PyTorch implementation of "Paralinguistics-Aware Speech-Empowered LLMs for Natural Conversation" (NeurIPS 2024)☆95Dec 3, 2024Updated last year
- A method that directly addresses the modality gap by aligning speech token with the corresponding text transcription during the tokenizat…☆119Sep 3, 2025Updated 10 months ago
- A family of state-of-the-art Transformer-based audio codecs for low-bitrate high-quality audio coding.☆437Jul 17, 2026Updated last week
- Adaptive Multimodal Reasoning via Reinforcement Learning☆23Jan 11, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Audio-FLAN☆161Sep 23, 2025Updated 10 months ago
- A curated list of Vision (video/image) to Audio Generation☆107Feb 10, 2026Updated 5 months ago
- Versatile Evaluation of Speech and Audio☆424Updated this week
- small audio language model for reasoning☆88Dec 4, 2025Updated 7 months ago
- [ICASSP 2024] TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models☆187Nov 22, 2024Updated last year
- [ICLR 2026] AudioMCQ: A 571k audio multiple-choice question dataset for post-training Large Audio Language Models with dual CoT annotatio…☆51Apr 21, 2026Updated 3 months ago
- ☆73Jul 17, 2026Updated last week