Benchmarking STT service TTFB and semantic WER for real-time AI applications
☆121Sep 18, 2026Updated this week
Alternatives and similar repositories for stt-benchmark
Users that are interested in stt-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A long-context eval☆161Aug 22, 2026Updated last month
- A lightweight library for normalizing speech transcripts before computing WER☆29Updated this week
- plugin manager for OpenVoiceOS , STT/TTS/Wakewords that can be used anywhere☆14Updated this week
- A Medical / Clinical Note Taking Demo Application using Deepgram Voice Agent API☆19Jul 9, 2025Updated last year
- Text-to-Speech Benchmark☆30Aug 18, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Skill library for building Pipecat bots with Claude Code☆26Jul 13, 2026Updated 2 months ago
- Train no-reference speech quality estimators with multiple datasets via learned, per-dataset alignments.☆19Aug 1, 2025Updated last year
- A real-time and multilingual speech translation model☆273Sep 9, 2026Updated 2 weeks ago
- Examples and Demos using the Cohere APIs☆23Nov 3, 2023Updated 2 years ago
- GUI Tool to create, manage and test Keyword Spotting models using TF 2.0☆13Feb 1, 2021Updated 5 years ago
- Revisiting ASR in the Age of Voice Agents [COLM26]☆30Apr 13, 2026Updated 5 months ago
- SLMTokBench for paper "SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models"☆37Aug 29, 2023Updated 3 years ago
- Meanflow and multilingual for F5-TTS model☆16Aug 23, 2025Updated last year
- HunyuanDiT with TensorRT and libtorch☆18May 22, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This is a fork of the original fairseq repository (version 0.12.2) with added classes for training mHuBERT-147.☆21Nov 19, 2024Updated last year
- A low-level Pipecat debugger.☆133Updated this week
- M#! Distributed shell pipelines with GNU Guile.☆14Dec 28, 2020Updated 5 years ago
- Open Source AI Benchmarking toolkit for benchmarking speech to text services☆59Apr 17, 2024Updated 2 years ago
- ☆61Jul 23, 2026Updated 2 months ago
- Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSI…☆25Oct 8, 2025Updated 11 months ago
- Text Normalization utilities for normalizing text for TTS☆26Mar 4, 2026Updated 6 months ago
- Implementation of the paper "Variable Bitrate Residual Vector Quantization for Audio Coding"☆11Apr 10, 2025Updated last year
- ☆63Apr 1, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Survey of available speech datasets for Polish ASR development☆17Jan 1, 2025Updated last year
- A baseline Automatic Speech Recognition system for Polish based on Kaldi.☆18Dec 21, 2021Updated 4 years ago
- Many ASRs under one roof. With Benchmarking... answering the question. What is the best ASR for my dataset?☆19Oct 5, 2022Updated 3 years ago
- Soniox Compare. Compare real-time voice AI side by side. No glossy charts, just results.☆38Sep 14, 2026Updated last week
- GNU Guile Matrix network SDK.☆14Nov 16, 2022Updated 3 years ago
- ☆25Apr 2, 2026Updated 5 months ago
- SDK for Daily's Video Component System (VCS)☆31Aug 11, 2026Updated last month
- Sample implementation of SRFI 197, a port of Clojure threading macros☆10Sep 6, 2020Updated 6 years ago
- An official implementation of Style-Talker for Spoken Dialogue Generation☆23Jan 12, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Gemini Multimodal Live command line client☆13Oct 14, 2025Updated 11 months ago
- A Slate plugin to handle onChange event on silence without event stack. Useful for implementing auto save Editor.☆15Jul 17, 2023Updated 3 years ago
- Reproducible voice-AI benchmarking — TTS / STT / S2S latency and accuracy.☆63Updated this week
- A universal phone recognizer that can transcribe speech in 70+ languages into IPA☆42Updated this week
- The implement of LLMTreeRec☆14Dec 9, 2024Updated last year
- A Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS☆65Dec 11, 2024Updated last year
- A Comprehensive Speech Processing Algorithms Library for research and production use☆18Oct 25, 2025Updated 10 months ago