A repository used to organize content related to Large Speech(Audio) Model, including paper, data, applications, tools and so on.
☆28Nov 8, 2025Updated 10 months ago
Alternatives and similar repositories for Awesome-Large-Speech-Model
Users that are interested in Awesome-Large-Speech-Model are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- S3PRL for Speech Emotion Recognition (see s3prl > downstream)☆15Feb 28, 2026Updated 6 months ago
- Building a inclusive, scalable, and high-performance multilingual translation model☆128May 7, 2026Updated 4 months ago
- We present a list of languages with their codes, families, regions and etc. We also present a list of multi-lingual corpora (with urls).☆87Jun 2, 2021Updated 5 years ago
- Official Code for "Learning to Reason via Mixture-of-Thought for Logical Reasoning"☆33Nov 20, 2025Updated 10 months ago
- A list of conferences and journals relevant to machine translation☆33Mar 17, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vi…☆124Jun 18, 2025Updated last year
- [IEEE, TASLP, 2023] The code of the paper "Multi-Source Discriminant Subspace Alignment for Cross-Domain Speech Emotion Recognition".☆19Sep 27, 2024Updated last year
- [AAAI 2026 & ACL 2026] The official implementation of the DIFFA series for dLLM-based large audio language model☆84Apr 7, 2026Updated 5 months ago
- Hybrid f0 estimation using Convolutional Neural Network☆12Apr 29, 2019Updated 7 years ago
- The implementation of Paper: Compose Yourself: Average-Velocity Flow Matching for One-Step Speech Enhancement.☆21Aug 1, 2026Updated last month
- An introduction to basic concepts of Transformers and key techniques of their recent advances.☆53Dec 21, 2023Updated 2 years ago
- The code for AAAI 2025 “Large Language Models Are Read/Write Policy-Makers for Simultaneous Generation”☆15Jan 3, 2025Updated last year
- CoNeTTE: An efficient Audio Captioning system leveraging multiple datasets with Task Embedding☆23Dec 17, 2025Updated 9 months ago
- [ICASSP 2024] KNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels☆42Mar 20, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICANN 2023] Anomaly-Based Insider Threat Detection via Hierarchical Information Fusion☆18Nov 20, 2023Updated 2 years ago
- A survey of rubrics across the evolving LLM landscape.☆70Jul 2, 2026Updated 2 months ago
- fastNLP reimplementation of the paper "A Novel Cascade Binary Tagging Framework for Relational Triple Extraction"☆11Dec 11, 2020Updated 5 years ago
- Advancing Block Diffusion Language Models for Test-Time Scaling☆16Feb 14, 2026Updated 7 months ago
- SpeechBrain中文文档☆12Mar 20, 2021Updated 5 years ago
- Repo for the FB AI Speech team.☆27Aug 24, 2021Updated 5 years ago
- ☆16Jun 22, 2017Updated 9 years ago
- ☆11Nov 16, 2024Updated last year
- Paper List☆18Jul 2, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Open-Source Turn-Taking Detection Model and Dataset for Full-Duplex Spoken Dialogue Systems☆145Jan 25, 2026Updated 7 months ago
- Jupyter notebooks for the book "Deep Learning with Python"☆11Aug 24, 2020Updated 6 years ago
- [NAACL 2024] Better Zero-Shot Reasoning with Role-Play Prompting☆36Nov 14, 2023Updated 2 years ago
- Official Implementation of "Prefix tuning for Automated Audio Captioning(ICASSP 2023)"☆30Dec 6, 2023Updated 2 years ago
- ☆18Mar 27, 2020Updated 6 years ago
- 세종말뭉치 가공데이터 Repository☆14Sep 11, 2018Updated 8 years ago
- ☆15Nov 26, 2024Updated last year
- A demo to show how to convert a TensorFlow model to TensorRT uff or PLAN☆11Jul 22, 2018Updated 8 years ago
- speech-dereverberation-using-GANs☆13Jan 28, 2019Updated 7 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- mnn asr demo.☆27Mar 24, 2025Updated last year
- Testing github connection on vscode update☆17Jun 1, 2026Updated 3 months ago
- 以Word2Vec和LSTM为基础,实现一个语言模型☆11Nov 7, 2017Updated 8 years ago
- [DATE 2023] Pipe-BD: Pipelined Parallel Blockwise Distillation☆12Jul 13, 2023Updated 3 years ago
- ☆29Apr 22, 2024Updated 2 years ago
- Zero-shot Domain-sensitive Speech Recognition with Prompt-conditioning Fine-tuning (ASRU2023)☆26Oct 10, 2023Updated 2 years ago
- Code for SLT 2016 paper on Grapheme-to-Phoneme conversion using attention based encoder-decoder models☆15Feb 20, 2019Updated 7 years ago