Offline streaming speech-to-text in the browser
☆25Aug 28, 2025Updated 10 months ago
Alternatives and similar repositories for wasm-speech-streaming
Users that are interested in wasm-speech-streaming are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆13Oct 11, 2024Updated last year
- Text-to-text alignment algorithm for speech recognition error analysis.☆31Jun 23, 2026Updated 3 weeks ago
- ☆30Apr 29, 2026Updated 2 months ago
- Open-weights voice acting pipeline combining zero-shot voice cloning with natural-language direction. Provide a reference voice (or gener…☆16May 25, 2026Updated last month
- ☆17Dec 18, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Parallelized automatic corpus collection for ASR. Forked from https://github.com/EgorLakomkin/KTSpeechCrawler☆23Mar 21, 2021Updated 5 years ago
- Forced alignment decoder for Whisper.☆16Mar 13, 2024Updated 2 years ago
- ☆19Mar 22, 2024Updated 2 years ago
- ☆20Sep 2, 2024Updated last year
- Train no-reference speech quality estimators with multiple datasets via learned, per-dataset alignments.☆18Aug 1, 2025Updated 11 months ago
- High-performance, semantic turn detection for conversational AI☆44Oct 1, 2025Updated 9 months ago
- Implementation of the paper "Variable Bitrate Residual Vector Quantization for Audio Coding"☆11Apr 10, 2025Updated last year
- 🎵 muse: Music Separation☆11Feb 14, 2024Updated 2 years ago
- T5Voice is a lightweight PyTorch implementation of T5-based text-to-speech synthesis, supporting both streaming and non-streaming speech …☆28Nov 7, 2025Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Target speaker automatic speech recognition (TS-ASR)☆14Oct 14, 2023Updated 2 years ago
- This repository presents an evaluation framework for speech-to-speech (S2S) models, following the methodology described in the EmphAsses …☆25Jan 9, 2024Updated 2 years ago
- Este repositorio contiene el dataset para el entrenamiento de una CNN que clasifica 10 diferentes tipos de especies de aves☆13Oct 8, 2021Updated 4 years ago
- Speaker-aware CTC (SACTC) for multi-talker overlapped speech recognition.☆22May 26, 2025Updated last year
- Collection of scripts from mHuBERT-147.☆35Nov 19, 2024Updated last year
- ☆12Mar 11, 2025Updated last year
- ☆13Sep 25, 2024Updated last year
- ☆34Jun 15, 2021Updated 5 years ago
- ☆12Jul 26, 2021Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆28Mar 11, 2026Updated 4 months ago
- Identifying the language of input text using character-level n-grams, with support for 45 languages☆11Dec 26, 2022Updated 3 years ago
- Speaker embedding for anime speech domain based on ECAPA_TDNN☆21Jun 22, 2025Updated last year
- ESLTTS dataset☆16Feb 6, 2025Updated last year
- A minimalist Docker project to help people getting started with Node, WizardCoder, CTransformers, Python, Express and TypeScript. Ready t…☆14Jun 23, 2023Updated 3 years ago
- StyleTTS 2 Optimized Training Fork☆32Feb 2, 2025Updated last year
- Repository for speech paper reading☆33Aug 19, 2021Updated 4 years ago
- ☆33Feb 4, 2025Updated last year
- ☆14Oct 3, 2025Updated 9 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [TASLP 2024] Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation☆31Sep 6, 2024Updated last year
- pytorch model for contexless-phoneme prediction from speech audio☆32Oct 30, 2025Updated 8 months ago
- Dlib's face recognition model evaluation on LFW dataset☆13Oct 3, 2017Updated 8 years ago
- Official implementation of the paper "Distilling a Pretrained Language Model to a Multilingual ASR Model" (Interspeech 2022)☆12Mar 12, 2024Updated 2 years ago
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆16Jun 16, 2024Updated 2 years ago
- Inference code for Interspeech 2025 paper, "LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec"☆36Oct 23, 2025Updated 8 months ago
- Fully local voice interface for Claude Code on Apple Silicon. Parakeet STT + Kokoro TTS + SmartTurn EOU + dual VAD.☆30Mar 24, 2026Updated 3 months ago