Real-time Voice Activity Detection (VAD) with some example use case like simple voice bot and live transcription (realtime transcription)
☆113Aug 18, 2025Updated last year
Alternatives and similar repositories for voice-activity-detection-vad-realtime
Users that are interested in voice-activity-detection-vad-realtime are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- On-device voice activity detection (VAD) powered by deep learning☆270Updated this week
- Welcome to the Real-Time Voice Activity Detection (VAD) program, powered by Silero-VAD model! 🚀 This program allows you to perform live …☆13Jul 9, 2023Updated 3 years ago
- Silero VAD: pre-trained enterprise-grade Voice Activity Detector☆10,251Updated this week
- Screenshots in record time - up to 2.5x faster than MSS (Multiple Screen Shots)☆12May 19, 2023Updated 3 years ago
- This is a single-speaker neural text-to-speech (TTS) system capable of training in a end-to-end fashion. It is inspired by the Tacotron a…☆12Dec 28, 2018Updated 7 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Community-controlled voice data collection for language preservation and AI development. Companion to 'AI Techniques for Indigenous Cultu…☆72May 6, 2026Updated 4 months ago
- Implementation of F5-TTS in MLX☆14Dec 13, 2024Updated last year
- Datasets for turn-taking research☆22Dec 21, 2023Updated 2 years ago
- bvh_broadcaster: broadcasting bvh motion capture as tf in ROS☆11Nov 26, 2018Updated 7 years ago
- Develop speaker recognition model based on i-vector using TIMIT database☆16Jul 4, 2019Updated 7 years ago
- Building a multi-agent RAG system with advanced RAG methods☆13Jan 12, 2025Updated last year
- The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained mode…☆13Jul 30, 2024Updated 2 years ago
- Whisper combined with Silero VAD, for improved long-form transcriptions☆55Dec 11, 2022Updated 3 years ago
- Linux SerialPort 串口Demo☆13Jan 26, 2017Updated 9 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆14Jan 4, 2025Updated last year
- 🏆 Ambassador Paper for Innovative Use of NLP for Building Educational Applications 2023: Is ChatGPT a Good Teacher Coach? Measuring Zero…☆15Jul 21, 2024Updated 2 years ago
- ☆14Jul 28, 2023Updated 3 years ago
- Agent-WebVoyager autonomously navigates the web like a human, performing tasks without specific APIs. It uses visual cues and intelligent…☆14Feb 13, 2024Updated 2 years ago
- A study about the Generalized Cross-Correlation with Phase Transform algorithm.☆14Nov 23, 2021Updated 4 years ago
- Deploy your GGML models to HuggingFace Spaces with Docker and gradio☆37Jun 6, 2023Updated 3 years ago
- This repo contains the baseline model recipes and pre-trained model for GramVanni hindi ASR challenge☆16Mar 26, 2022Updated 4 years ago
- Package for linking messages from vive to pose in moveit☆11Jan 2, 2021Updated 5 years ago
- A python library / model for creating co-references between AMR graph nodes.☆11Dec 11, 2022Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [WIP] Layer Diffusion for WebUI (via Forge)☆13Apr 21, 2024Updated 2 years ago
- Target Speaker Extraction Toolkit☆314Oct 4, 2025Updated 11 months ago
- Self-Supervised Speech/Sound Pre-training and Representation Learning Toolkit☆13Nov 18, 2022Updated 3 years ago
- A mini, simple, and fast end-to-end automatic speech recognition toolkit.☆53Dec 6, 2022Updated 3 years ago
- Embeddable ringcentral phone for hubspot(Google Chrome extension)☆13Mar 4, 2023Updated 3 years ago
- PyTorch implementation of STAGE model☆16Mar 17, 2025Updated last year
- Bilingual Singing Voice Synthesis☆18Mar 25, 2024Updated 2 years ago
- Speech recognition with ESP32 and Edge Impulse.☆12Apr 25, 2024Updated 2 years ago
- A multi-agent business consultant app on streamlit implemented using crewAI☆19Jul 5, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Choose your own adventure with LLMs☆23May 27, 2025Updated last year
- Open Source Text Embedding Models with OpenAI Compatible API☆171Jul 13, 2024Updated 2 years ago
- Deconstructing the black box that is the algorithm underlying the sleep score provided by Fitbit☆13Aug 6, 2020Updated 6 years ago
- Implementation of Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt (NAACL'24).☆119Jan 26, 2025Updated last year
- Official implementation of the ECCV2024 paper: Generalizable Facial Expression Recognition☆22Sep 20, 2024Updated 2 years ago
- ☆18May 27, 2024Updated 2 years ago
- ☆11Oct 27, 2019Updated 6 years ago