stt websockect server using sherpa-onnx
☆57Feb 28, 2026Updated 4 months ago
Alternatives and similar repositories for stt-server
Users that are interested in stt-server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- On-device speech AI runtime for ASR, TTS, VAD, and voice cloning. Python-simple, C++-native, GGUF-powered.☆22Jul 15, 2026Updated last week
- Port of Funasr's Paraformer model in C/C++☆43Jun 19, 2024Updated 2 years ago
- 一个基于 Sherpa-ONNX 的高性能语音识别服务,支持实时VAD(语音活动检测)、多语言语音识别和声纹识别功能。☆115Jan 4, 2026Updated 6 months ago
- SummerTTS 是一个基于C++的独立编译的中文和英文语音合成项目,可以本地运行不需要网络,而且没有额外的依赖,一键编译完成即可用于中文和英文的语音合成。SummerTTS is a standalone Chinese and English speech synt…☆25Aug 17, 2024Updated last year
- ☆25Mar 8, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Streaming ASR and TTS based on FastAPI+ sherpa-onnx☆222Nov 2, 2025Updated 8 months ago
- Port of Funasr's Sense-voice model in C/C++☆568Dec 19, 2025Updated 7 months ago
- Moss: A voice assistant using LLM and Langchain which can control your home assistant and chat more.☆32Jul 12, 2025Updated last year
- ☆15Dec 22, 2025Updated 7 months ago
- FreeSWITCH Module for MP3 recording☆11Feb 23, 2020Updated 6 years ago
- Offline Speaker Diarization with SenseVoice by Sherpa ONNX.☆15Dec 23, 2024Updated last year
- A lightweight demo of FunASR-Nano using ONNX runtime.☆83Feb 25, 2026Updated 5 months ago
- High-performance Qwen3-TTS implementation | Instruction-driven · Zero-shot voice cloning · Streaming · RTF 0.55☆66Jun 26, 2026Updated 3 weeks ago
- low-latency realtime ASR based on FireRedASR☆62Jul 8, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Mar 30, 2023Updated 3 years ago
- GTCRN(ncnn).☆19May 22, 2025Updated last year
- Hacked FreeSWITCH-G729 speech codec using Intel® Integrated Performance Primitive.☆23Jan 22, 2013Updated 13 years ago
- LLM-based ASR recipe with Zipformer encoder and Qwen LLM☆35Sep 25, 2025Updated 10 months ago
- ☆56Sep 17, 2025Updated 10 months ago
- (WIP)long form speech generatoins☆30Apr 2, 2025Updated last year
- A playground for experimenting with acoustic echo cancellation using a microphone, speaker, and ONNX.☆13Oct 22, 2024Updated last year
- A enterprise-grade Voice Activity Detector from modelscope and funasr.☆139Apr 26, 2023Updated 3 years ago
- The CPP version of Silero VAD: pre-trained enterprise-grade Voice Activity Detector☆23May 11, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- TEN VAD low-latency voice activity detection for real-time streaming, integrated with livekit-agents☆26Nov 13, 2025Updated 8 months ago
- Pseudo Streaming SenseVoice with Hotwords☆466Jun 15, 2026Updated last month
- 一键将视频转换为优质小红书笔记,自动优化内容和配图;追加了可以读取本地视频的功能☆12Dec 22, 2024Updated last year
- A package used to test webrtc apm functions, such as aec, ns☆17Feb 21, 2019Updated 7 years ago
- A streaming audio reader, processor, and writer built on top of soundfile, and PyAV (bindings for FFmpeg)☆39Mar 31, 2026Updated 3 months ago
- Clean up noisy speech in real time with DPDFNet - open-source streaming speech enhancement for research, audio apps, and edge devices. In…☆113Updated this week
- An open source chat bot architecture for voice/vision (and multimodal) assistants, local(CPU/GPU bound) and remote(I/O bound) to run.☆89Dec 28, 2025Updated 6 months ago
- Template for creating audio encoders compatible with X-ARES☆19Feb 11, 2026Updated 5 months ago
- paraformer web server build with sanic☆28May 3, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- This repo is an exploratory experiment to enable frozen pretrained RWKV language models to accept speech modality input. We followed the …☆54Dec 23, 2024Updated last year
- We Speech Transcript based on LLM, in 300 lines of code.☆182Jun 20, 2025Updated last year
- Dynamic Mixing For Speech Processing (mix-on-the-fly)☆22Jul 19, 2022Updated 4 years ago
- Freeswitch ASR module☆22Jan 15, 2025Updated last year
- Causal streaming adaptation of OpenAI Whisper for real-time transcription on small audio chunks.☆75Mar 31, 2026Updated 3 months ago
- ☆21Nov 28, 2025Updated 7 months ago
- X-ASR is a series of automatic speech recognition models based on the icefall framework, focusing on streaming ASR and low-latency deploy…☆145Jul 8, 2026Updated 2 weeks ago