☆143Jan 24, 2026Updated 7 months ago
Alternatives and similar repositories for Speech-and-audio-papers-Top-Conference
Users that are interested in Speech-and-audio-papers-Top-Conference are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PyTorch implementation of USR 2.0 (ICLR 2026)☆16Sep 9, 2026Updated 2 weeks ago
- Audio Codec Speech processing Universal PERformance Benchmark☆311Jul 4, 2026Updated 2 months ago
- 🦇 Encoder of BAT (Learning to Reason about Spatial Sounds with Large Language Models)☆90Feb 13, 2025Updated last year
- Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with syntheti…☆137Mar 3, 2026Updated 6 months ago
- Awesome speech/audio LLMs, representation learning, and codec models☆1,248Jul 10, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- (ICLR 2025) Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation☆16Apr 29, 2025Updated last year
- Update ASR paper everyday☆514May 16, 2026Updated 4 months ago
- Understanding and Tackling Hallucinations in Large Audio-Language Models | ICASSP 2025, Interspeech 2024☆34Mar 14, 2025Updated last year
- Awesome Neural Codec Models, Text-to-Speech Synthesizers & Speech Language Models☆248Jul 9, 2026Updated 2 months ago
- ☆46Apr 2, 2025Updated last year
- [ACL 2025 Main] UniCodec: a unified audio codec with a single codebook to support multi-domain audio data, including speech, music, and s…☆158May 30, 2025Updated last year
- The implementation for "Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions"☆51Apr 7, 2025Updated last year
- VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling☆100Nov 9, 2024Updated last year
- Unofficial pytorch reproduction for the paper "Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction" (…☆60Apr 4, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)☆668Updated this week
- The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio gener…☆958Updated this week
- Evaluation Protocol for Large-Scale Zero-Shot TTS Literature☆98Mar 12, 2025Updated last year
- [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix☆222Feb 25, 2026Updated 6 months ago
- [ACL 2026 Main] MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows☆150Sep 2, 2025Updated last year
- Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'☆166Mar 26, 2026Updated 5 months ago
- LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement☆107Apr 1, 2025Updated last year
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling☆177Feb 28, 2025Updated last year
- [ICCV'25] Official PyTorch Implementation of "VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models"☆17Dec 8, 2025Updated 9 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [CVPR 2024] AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation☆48Sep 6, 2024Updated 2 years ago
- [NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”☆38Jul 24, 2025Updated last year
- This is an evolving repo for the paper "Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey".☆286Updated this week
- [ACM Multimedia 24]The official repository of SpeechCraft dataset, a large-scale expressive bilingual speech dataset with natural languag…☆202Feb 28, 2026Updated 6 months ago
- Curated list for papers, codes and resources related to Text-to-Audio (TTA) Generation☆77Jul 20, 2026Updated 2 months ago
- Official PyTorch implementation of "Paralinguistics-Aware Speech-Empowered LLMs for Natural Conversation" (NeurIPS 2024)☆95Dec 3, 2024Updated last year
- A method that directly addresses the modality gap by aligning speech token with the corresponding text transcription during the tokenizat…☆121Sep 3, 2025Updated last year
- A family of state-of-the-art Transformer-based audio codecs for low-bitrate high-quality audio coding.☆445Jul 17, 2026Updated 2 months ago
- Adaptive Multimodal Reasoning via Reinforcement Learning☆24Jan 11, 2026Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Audio-FLAN☆163Sep 23, 2025Updated last year
- A curated list of Vision (video/image) to Audio Generation☆111Aug 24, 2026Updated 3 weeks ago
- Versatile Evaluation of Speech and Audio☆437Updated this week
- small audio language model for reasoning☆89Dec 4, 2025Updated 9 months ago
- [ICASSP 2024] TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models☆188Nov 22, 2024Updated last year
- [ICLR 2026] AudioMCQ: A 571k audio multiple-choice question dataset for post-training Large Audio Language Models with dual CoT annotatio…☆52Apr 21, 2026Updated 5 months ago
- ☆76Jul 29, 2026Updated last month