This is a repository for fine-tuning Qwen2-Audio, currently supporting Distributed Data Parallel (DDP) and DeepSpeed.
☆50Jul 28, 2025Updated last year
Alternatives and similar repositories for Qwen2-Audio-finetune
Users that are interested in Qwen2-Audio-finetune are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Colab notebook for fine-tuning Qwen2-Audio with trl's SFT and PPO trainers.☆24Nov 23, 2024Updated last year
- The project is associated with the recently-launched INTERSPEECH 2025 Workshop on Multilingual Conversational Speech Language Model (MLC-…☆51May 14, 2025Updated last year
- 🤗 R1-AQA Model: mispeech/r1-aqa☆325Mar 28, 2025Updated last year
- This is a framework for using large language models to improve ASR recognition accuracy. You need to provide the recognized text and tag …☆18Jun 5, 2025Updated last year
- ☆23Oct 17, 2024Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Simplistic Implementation of Zipformer:A faster and better encoder for automatic speech recognition in PyTorch☆22Jun 3, 2024Updated 2 years ago
- [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix☆214Feb 25, 2026Updated 5 months ago
- The repoduction codes for Qwen-Audio Fine-tuning☆55Feb 28, 2026Updated 5 months ago
- Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style☆15Aug 18, 2025Updated 11 months ago
- CTC decoder with hotwords for ASR.☆38Jun 15, 2026Updated last month
- FireRedChat pVAD plugin for LiveKit Agents☆22Sep 16, 2025Updated 10 months ago
- Target speaker automatic speech recognition (TS-ASR)☆14Oct 14, 2023Updated 2 years ago
- [CVPR 2025] Official implementation of paper "Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie…☆23Jun 6, 2025Updated last year
- TASU: A New Style of Alignment of Speech LLM with only Text Training Data, zero-shot on ASR and Other SU tasks☆27Jul 20, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆16Nov 4, 2025Updated 8 months ago
- Official implementation of the paper titled "Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Mu…☆28Mar 5, 2024Updated 2 years ago
- This repository is for the paper Incorporating External POS Tagger for Punctuation Restoration. Proc. Interspeech 2021, 1987-1991, doi: 1…☆11May 24, 2026Updated 2 months ago
- ☆116Oct 21, 2025Updated 9 months ago
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning (🔥The Exploration of R1 for General Audio-Vis…☆78Jun 3, 2026Updated last month
- One command to start a streaming ASR server.☆12Oct 2, 2024Updated last year
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness (https://arxiv.org/abs/2511.07931)☆79Dec 23, 2025Updated 7 months ago
- Code for the paper "JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis"☆14Nov 5, 2024Updated last year
- Code for DeSTA2.5-Audio, general-purpose LALM☆141Feb 4, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.☆142Apr 7, 2026Updated 3 months ago
- SLT 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge☆12Jun 11, 2024Updated 2 years ago
- The baseline system for the ICASSP2024 ICMC-ASR Challenge.☆57Dec 6, 2023Updated 2 years ago
- Open, royalty free, lyrics2song / song generation data collection / cleaning pipeline.☆17May 9, 2025Updated last year
- Prompting Large Language Models with Audio for General-Purpose Speech Summarization☆20May 14, 2025Updated last year
- Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model☆25May 21, 2026Updated 2 months ago
- A Framework for Speech, Language, Audio, Music Processing with Large Language Model☆1,050Jan 15, 2026Updated 6 months ago
- The baselines of ARC-Challenge-Interspeech2026☆60Dec 1, 2025Updated 7 months ago
- ☆16Apr 2, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- sherpa with mlx☆15Aug 2, 2025Updated 11 months ago
- Acoustic echo cancelation(AEC) is a main algorithm in the pipe line of acoustic devices with KWS or ASR. FNLMS is used.☆19Apr 22, 2019Updated 7 years ago
- Dataset for Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models in Interspeech 2024.☆16Jul 4, 2024Updated 2 years ago
- UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts☆41Jun 12, 2025Updated last year
- real-time speech enhance☆18Jan 23, 2024Updated 2 years ago
- FCTalker: Fine and Coarse Grained Context Modeling for Expressive Conversational Speech Synthesis (Accepted by ISCSLP'2024)☆26Feb 22, 2024Updated 2 years ago
- A Massive Contextual Speech Recognition Benchmark.☆107Aug 6, 2025Updated 11 months ago