This UI serves as a Synthetic ASR Dataset Generator powered by/for OpenAI Whisper, enabling users to capture audio, transcribing it, on the fly and manage the generated dataset 🤗. Fine tune Whisper or enhanced and custom datasets
☆34Nov 26, 2024Updated last year
Alternatives and similar repositories for Whisper-Synthetic-ASR-Dataset-Generator
Users that are interested in Whisper-Synthetic-ASR-Dataset-Generator are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆10Apr 16, 2020Updated 6 years ago
- A Python package for converting numbers expressed in natural language to numerical values.☆13Nov 25, 2023Updated 2 years ago
- Rust bindings for the Voikko library☆16Mar 26, 2022Updated 4 years ago
- ☆34May 15, 2023Updated 3 years ago
- A web application that predicts the nationality of a person's name☆10Dec 23, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Can Neural Networks reconstruct missing audio data? What about GANs?☆18Nov 6, 2019Updated 6 years ago
- Analyze and visualize how rhythm, timbre, loudness, pitch, spectral characteristics and other key audio features evolve over time across …☆11Jul 27, 2026Updated 3 weeks ago
- Speech-to-text transcription VST3/ARA plugin☆63Jun 8, 2026Updated 2 months ago
- A lightweight header-only c++ library for real time audio applications, oriented to the embedded world.☆18Jul 23, 2021Updated 5 years ago
- Demo how to use use Unix Domain Sockets in Swift on macOS.☆13Sep 24, 2021Updated 4 years ago
- ComfyUI port of SDWebUI Vectorscope CC and Diffusion CG extensions☆21Feb 24, 2025Updated last year
- Custom node for ComfyUI. Add a node for drawing text to the area of SEGS.☆14Mar 30, 2025Updated last year
- Voice activity detection and speaker gender segmentation audiovisual corpus☆16Jan 20, 2025Updated last year
- Single Image Haze Removal Using AODNet in Pytorch☆15Mar 5, 2021Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A Java port of whisper 3, based on the huggingface version, using DJL.☆17Apr 3, 2024Updated 2 years ago
- Comparing Audio Features for Unsupervised Sound Classification☆10Jun 22, 2022Updated 4 years ago
- Trainer and Evaluation scripts for fine-tuning Whisper models for the Ukrainian language☆23Jan 13, 2023Updated 3 years ago
- ☆10Aug 3, 2019Updated 7 years ago
- ☆12May 1, 2019Updated 7 years ago
- Official source code for the paper "Tailored Design of Audio-Visual Speech Recognition Models using Branchformers"☆15Feb 24, 2025Updated last year
- Scripts to convert audio files to spectrograms and back☆12Nov 23, 2017Updated 8 years ago
- 一个强调工程化、可观测、可测试、可扩展的 RAG 项目。TraceRAG 的目标不是只把答案“生成出来”,而是把文档导入、切块、向量化、检索、带来源回答、评估与后续 tracing 拆成可独立验证的阶段,逐步演进成一个可维护、可解释、可复盘的生产级 RAG。☆15Apr 2, 2026Updated 4 months ago
- ARCHIVED Fetch 'Scholary' Full Text from 'Crossref'☆17May 10, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Aug 25, 2021Updated 4 years ago
- Multi-lingual AudioCaps☆14Nov 20, 2023Updated 2 years ago
- Code for InterSpeech 2024 Paper: LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition☆19Jul 16, 2024Updated 2 years ago
- Hardware and support board schematics☆17Nov 10, 2016Updated 9 years ago
- Collaborative audio annotation tool☆17Sep 16, 2022Updated 3 years ago
- VHDL package for reading formatted data from comma-separated-values (CSV) files☆23Sep 10, 2013Updated 12 years ago
- Encode an image to sound (WAV file) and view it as a spectrogram. Optimized Python 3 version.☆11Jan 25, 2023Updated 3 years ago
- Rainbowgram with Python☆13Jan 28, 2019Updated 7 years ago
- ☆19May 9, 2019Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Dimensionality reduction (UMAP, t-SNE, PCA) for ImageJ/Fiji☆12May 6, 2025Updated last year
- ☆28Jun 28, 2024Updated 2 years ago
- Keras implementation of conditional waveGAN. Application to knocking sound effects with emotion.☆10Jun 22, 2020Updated 6 years ago
- 抓取midjourney频道生产的图片☆11Jun 19, 2023Updated 3 years ago
- Convert images to audio for display in a spectrogram☆13Apr 17, 2018Updated 8 years ago
- CNN-to-FPGA-framework for small CNN, written in VHDL and Python☆24Jun 8, 2021Updated 5 years ago
- Using Deep Learning for singing voice separation - Project for the course DT2119 Speech and Speaker Recognition offered by KTH in 2018☆15Jun 16, 2018Updated 8 years ago