ina-foss/InaGVAD

Readme badge preview -

If you own this repo, copy the snippet below and add it to your README.md

[![RelatedRepos](https://img.shields.io/badge/related-repos-yellow)](https://relatedrepos.com/gh/ina-foss/InaGVAD)

ina-foss / InaGVAD

Voice activity detection and speaker gender segmentation audiovisual corpus

☆16

Alternatives and similar repositories for InaGVAD

Users that are interested in InaGVAD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

leto19 / WhiSQA
View on GitHub
Whisper Speech Quality Assessment (WhiSQA)
☆16Apr 14, 2026Updated 3 months ago
thevoicecompany / gazelle-train
View on GitHub
Joint speech-language model - respond directly to audio!
☆30May 13, 2024Updated 2 years ago
llm-lab-org / CLASP
View on GitHub
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
☆13Jun 27, 2025Updated last year
lattice-8094 / propp
View on GitHub
PROPP: A Python library for narrative analysis
☆25Jun 1, 2026Updated last month
iliassarbout / CityOfLight
View on GitHub
City of Light (COL) is a geospatially faithful, Unity-based digital twin of Paris enabling high-performance embodied simulation for AI an…
☆50Mar 31, 2026Updated 3 months ago
GPUs on demand by Runpod - Special Offer Available • Ad
Run AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
Aratako / MioCodec
View on GitHub
☆27Feb 14, 2026Updated 5 months ago
JethroWangSir / SincQDR-VAD
View on GitHub
☆26Aug 29, 2025Updated 10 months ago
JacobLinCool / MPSENet
View on GitHub
Python package of MP-SENet from Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement.
☆22Nov 1, 2024Updated last year
yfyeung / DS-WED
View on GitHub
[ICASSP 2026] Official code for "Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration"
☆17Apr 16, 2026Updated 3 months ago
pilot7747 / VoxDIY
View on GitHub
This repository provides data and code for "Vox Populi, Vox DIY: Benchmark Dataset for Crowdsourced Audio Transcription" paper.
☆16Jul 22, 2021Updated 5 years ago
facebookresearch / spidr
View on GitHub
This repository contains the training code from paper "SpidR Learning Fast and Stable Linguistic Units for Spoken Language Models Without…
☆57Updated this week
lcn-kul / xls-r-analysis-sqa
View on GitHub
Analysis of XLS-R for Speech Quality Assessment
☆15Feb 10, 2025Updated last year
NTIA / WEnets
View on GitHub
Reference Implementations of Waveform Evaluation Networks (WEnets)
☆27Sep 18, 2023Updated 2 years ago
Okrio / FSPEN
View on GitHub
☆21Apr 27, 2024Updated 2 years ago
End-to-end encrypted email - Proton Mail • Ad
Special offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
Yifei-ZHAO96 / Tr-VAD
View on GitHub
Tr-VAD: An Efficient Transformer based Voice Activity Detection Model
☆18Aug 1, 2024Updated last year
anindex / noprop
View on GitHub
☆13May 9, 2025Updated last year
laavanyebahl / OCR-Extracting-text-from-images-with-neural-networks
View on GitHub
OCR using a simple network developed from scratch on NIST36 dataset vs with CNN on PyTorch on EMNIST dataset
☆12Jan 29, 2019Updated 7 years ago
uthree / ddsp-vocoder
View on GitHub
☆12Nov 7, 2024Updated last year
JaesungHuh / av-diarization
View on GitHub
Audio-visual diarization pipeline used for creating VoxConverse dataset
☆22Jun 6, 2025Updated last year
gweltou / anaouder-cli
View on GitHub
Anaouder mouezh e Brezhoneg gant Vosk
☆15Nov 24, 2025Updated 8 months ago
b-sigpro / sed-hsmm
View on GitHub
Onset-and-Offset-Aware Sound Event Detection
☆21Feb 10, 2025Updated last year
kjhealy / covid_symptoms
View on GitHub
☆10Apr 16, 2020Updated 6 years ago
liuhuang31 / HiFTNet-sr
View on GitHub
HiFTNet wav/audio super-resolution 16/24 kHz to 48 kHz
☆24Jan 2, 2024Updated 2 years ago
Wordpress hosting with auto-scaling - Free Trial Offer • Ad
Fully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
nii-yamagishilab / speaker_sex_attribute_privacy
View on GitHub
Project for HIDING SPEAKER’S SEX IN SPEECH USING ZERO-EVIDENCE SPEAKER REPRESENTATION IN AN ANALYSIS/SYNTHESIS PIPELINE
☆15Nov 30, 2022Updated 3 years ago
Deep-unlearning / nano-cohere-transcribe
View on GitHub
Pure-PyTorch inference for CohereLabs/cohere-transcribe-03-2026 (2B Conformer + Transformer ASR, 14 languages).
☆37Apr 29, 2026Updated 2 months ago
ZXHY-82 / w2v-BERT-2.0_SV
View on GitHub
☆53Mar 28, 2026Updated 3 months ago
Many0therFunctions / MaskGCT-Text-To-Semantic-Finetune
View on GitHub
This is not remotely close to a finished product, and does not intend to nor does this claim to be working fine-tuning code for MaskGCT. …
☆13Dec 4, 2024Updated last year
bekirbakar / replay-attack-detection
View on GitHub
Deep learning-based audio spoofing attack detection experiments for speaker verification.
☆14Apr 20, 2023Updated 3 years ago
offerzen / wombat
View on GitHub
Optimises Webflow Core Web Vitals and integrates into Git by scraping and rehosting projects
☆17Jun 18, 2024Updated 2 years ago
hlt-mt / Speech-MASSIVE
View on GitHub
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSI…
☆25Oct 8, 2025Updated 9 months ago
ronjakoi / voikko-rs
View on GitHub
Rust bindings for the Voikko library
☆16Mar 26, 2022Updated 4 years ago
daanzu / kaldi_ag_training
View on GitHub
Docker image and scripts for training finetuned or completely personal Kaldi speech models. Particularly for use with kaldi-active-gramma…
☆21Jan 24, 2022Updated 4 years ago
Managed hosting for WordPress and PHP on Cloudways • Ad
Managed hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
angeluriot / 2D_fluid_simulation
View on GitHub
A program simulating a fluid in 2D.
☆15May 16, 2024Updated 2 years ago
tuanct1997 / Federated-Learning-ASR-based-on-wav2vec-2.0
View on GitHub
☆18Mar 13, 2024Updated 2 years ago
IDRnD / VoxTube
View on GitHub
The VoxTube dataset official repository
☆71Feb 14, 2024Updated 2 years ago
EKarton / Names-To-Nationality-Predicter
View on GitHub
A web application that predicts the nationality of a person's name
☆10Dec 23, 2024Updated last year
jhauret / vibravox
View on GitHub
Speech to Phoneme, Bandwidth Extension and Speaker Verification using the Vibravox dataset.
☆51Dec 1, 2025Updated 7 months ago
lehidalgo / mastering-llms
View on GitHub
In this repo you will find a roadmap to master Large Language Models
☆18Jun 12, 2023Updated 3 years ago
line / WaveTrainerFit
View on GitHub
Official implementation of "Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech G…
☆16Feb 6, 2026Updated 5 months ago