Acoustic Neighbor Embeddings
☆33Jul 13, 2025Updated last year
Alternatives and similar repositories for ml-acn-embed
Users that are interested in ml-acn-embed are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆10Jun 11, 2024Updated 2 years ago
- ☆12Mar 11, 2025Updated last year
- Tidy Tunes is an easy-to-use pipeline for mining high-quality audio data for speech generation models. To do so, it chains multiple open …☆23May 19, 2026Updated 2 months ago
- ESLTTS dataset☆16Feb 6, 2025Updated last year
- Using YouTube to prepare a speech recognition dataset for any language☆10Mar 30, 2021Updated 5 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Transformer based ASR Engine.☆13Aug 23, 2021Updated 4 years ago
- ☆15Feb 6, 2026Updated 5 months ago
- CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval☆13Jun 27, 2025Updated last year
- ☆12Nov 7, 2024Updated last year
- ☆11Nov 5, 2021Updated 4 years ago
- A wrapper for Audeering's wav2vec-based dimensional speech emotion recognition☆22Aug 9, 2023Updated 2 years ago
- DUSTED: Spoken-Term Discovery using Discrete Speech Units☆17Oct 2, 2024Updated last year
- This is the official PyTorch implementation for the paper "Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence …☆16Sep 3, 2025Updated 10 months ago
- ☆15Nov 26, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Sep 19, 2022Updated 3 years ago
- CTC decoder with hotwords for ASR.☆38Jun 15, 2026Updated last month
- NISQA - Non-Intrusive Speech Quality and TTS Naturalness Assessment☆16Apr 13, 2022Updated 4 years ago
- KABooks is a tool to automate the process of creating datasets for training Text-To-Speech (TTS) and Speech-To-Text (STT) models. Using a…☆13Mar 24, 2023Updated 3 years ago
- ☆14Jun 16, 2023Updated 3 years ago
- The official implementation of DMEL the method presented in the paper "DMEL: The differentiable log-Mel spectrogram as a trainable layer …☆24Dec 21, 2024Updated last year
- Official implementation of TISDiSS, a scalable framework for discriminative source separation.☆16Oct 19, 2025Updated 9 months ago
- [ICASSP 2025] AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder☆14Mar 11, 2025Updated last year
- Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS (E2 TTS) in MLX☆29Oct 15, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICASSP 2024] Official code for FreGrad☆35May 13, 2024Updated 2 years ago
- A vector DB so easy, even your grandparents can build a RAG system 😁☆25Apr 1, 2026Updated 3 months ago
- MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts☆33Jan 15, 2026Updated 6 months ago
- Indonesian speech/phoneme recognizer powered by Kaldi 2.0 (lhotse, icefall, sherpa).☆16Jun 30, 2023Updated 3 years ago
- ☆21Sep 14, 2025Updated 10 months ago
- An open-source Kazakh Emotional Text-to-Speech Dataset☆36Aug 1, 2025Updated 11 months ago
- Unsupervised phone and word segmentation using dynamic programming on self-supervised VQ features.☆39May 5, 2026Updated 2 months ago
- Variable Bitrate Residual Vector Quantization for Audio Coding☆54May 1, 2025Updated last year
- Ultra-low-bitrate Speech Codec for Speech Language Modeling Applications☆92Dec 20, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆17Jul 16, 2020Updated 6 years ago
- Train a fiwGAN or ciwGAN model using your own training data☆14Oct 13, 2022Updated 3 years ago
- [Interspeech 2024] Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement☆43Jul 25, 2025Updated last year
- [SLT'24] Mamba-based Decoder-Only Approach for Speech Recognition☆19Dec 1, 2024Updated last year
- Attention-Enhanced Short-Time Wiener Solution for Acoustic Echo Cancellation☆31Nov 12, 2025Updated 8 months ago
- Eureka-Audio: A 1.7B lightweight audio–language model that matches 7B–30B models on ASR, audio understanding, and paralinguistic reasonin…☆40Apr 11, 2026Updated 3 months ago
- ☆16Nov 11, 2024Updated last year