Practical, Colab-friendly notebooks for fine-tuning and running audio AI models
☆425Jul 22, 2026Updated last month
Alternatives and similar repositories for smol-audio
Users that are interested in smol-audio are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Oct 22, 2025Updated 10 months ago
- VyvoTTS: LLM-Based Text-to-Speech Training Framework☆261Aug 9, 2026Updated 3 weeks ago
- A package for NeuCodec: a 50hz, 0.8kbps, 24kHz audio codec.☆164Jun 22, 2026Updated 2 months ago
- Collection of scripts from mHuBERT-147.☆35Nov 19, 2024Updated last year
- MOSS-Audio is an open-source foundation model for unified audio understanding, enabling speech, sound, music, captioning, QA, and reasoni…☆651Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆30Jul 3, 2026Updated last month
- Pure-PyTorch inference for CohereLabs/cohere-transcribe-03-2026 (2B Conformer + Transformer ASR, 14 languages).☆44Apr 29, 2026Updated 4 months ago
- X-Voice☆181Aug 5, 2026Updated 3 weeks ago
- Fast audio super resolution from 16khz to 48khz.☆218Jan 3, 2026Updated 7 months ago
- Unofficial fairseq-free PyTorch implementation of UTMOS (v1, 2022), matching the original system.☆35Jun 6, 2026Updated 2 months ago
- Building actual open source including dataset Multilingual TTS more than 150 languages with Voice Cloning.☆57Updated this week
- Hosts text-to-speech corpus and speech synthesizers for African languages.☆20May 31, 2023Updated 3 years ago
- Open-source text-to-speech model from KRAFTON trained exclusively on public speech data, with curated datasets and reproducible training …☆96May 21, 2026Updated 3 months ago
- Lean neural real-time acoustic echo cancellation with soft delay estimation - GGML and PyTorch inference☆208Jul 20, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆18Nov 19, 2025Updated 9 months ago
- Descript Audio Codec - VAE Variant (.dac-vae): High-Fidelity Audio Compression with Variational Autoencoder☆40Aug 30, 2025Updated last year
- ☆38Oct 10, 2025Updated 10 months ago
- dawproject-py – A Python repository with code for parsing, generating, and modifying DAWProject files, enabling seamless DAW interoperabi…☆66Feb 22, 2026Updated 6 months ago
- Pure-PyTorch Parakeet TDT inference☆52Mar 10, 2026Updated 5 months ago
- Liquid Audio - Speech-to-Speech audio models by Liquid AI☆567Jun 5, 2026Updated 2 months ago
- ☆11Aug 7, 2024Updated 2 years ago
- Application to convert .png files to filmstrip☆11May 21, 2025Updated last year
- Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation☆75Jun 16, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆40Apr 3, 2025Updated last year
- Swift wrapper for llama.cpp☆17Jun 15, 2026Updated 2 months ago
- A resource embedder for C++/CMake☆34Jul 20, 2026Updated last month
- Scripts to create speech corpora from open.bible☆13Jan 3, 2022Updated 4 years ago
- ☆374Aug 28, 2025Updated last year
- A high quality and fast TTS repository☆522Dec 22, 2025Updated 8 months ago
- NISQA - Non-Intrusive Speech Quality and TTS Naturalness Assessment☆16Apr 13, 2022Updated 4 years ago
- C++17 port of librosa with wasm and SPM package. Done with agents. YMMV.☆53Jun 23, 2026Updated 2 months ago
- Text-to-text alignment algorithm for speech recognition error analysis.☆33Jun 23, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICASSP 2026] TinyMU: A Compact Audio Language Model for Music Understanding☆39Apr 20, 2026Updated 4 months ago
- Inference, Fine Tuning and many more recipes with Gemma family of models☆304Updated this week
- Omnilingual ASR Open-Source Multilingual SpeechRecognition for 1600+ Languages☆2,902Dec 30, 2025Updated 8 months ago
- Llama-Mimi is a speech language model that uses a unified tokenizer (Mimi) and a single Transformer decoder (Llama) to jointly model sequ…☆31Sep 20, 2025Updated 11 months ago
- A notebooks based (soft) intro to modern TTS☆18Jun 8, 2025Updated last year
- CML-TTS: A Multilingual Dataset for Speech Synthesis☆36Jul 31, 2024Updated 2 years ago
- Generative Expressive Conversational Speech Synthesis (Accepted by MM'2024)☆78Nov 1, 2024Updated last year