This UI serves as a Synthetic ASR Dataset Generator powered by/for OpenAI Whisper, enabling users to capture audio, transcribing it, on the fly and manage the generated dataset 🤗. Fine tune Whisper or enhanced and custom datasets
☆36Nov 26, 2024Updated last year
Alternatives and similar repositories for Whisper-Synthetic-ASR-Dataset-Generator
Users that are interested in Whisper-Synthetic-ASR-Dataset-Generator are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Text-based media editing interface☆16Aug 9, 2017Updated 9 years ago
- A fast CPU-first video/audio transcriber for generating caption files with Whisper and CTranslate2, hosted on Hugging Face Spaces.☆11Updated this week
- A Python package for converting numbers expressed in natural language to numerical values.☆14Nov 25, 2023Updated 2 years ago
- A real time offline transcriber with gui, based on OpenAI whisper☆17Dec 25, 2025Updated 9 months ago
- ☆33May 15, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Can Neural Networks reconstruct missing audio data? What about GANs?☆18Nov 6, 2019Updated 6 years ago
- Analyze and visualize how rhythm, timbre, loudness, pitch, spectral characteristics and other key audio features evolve over time across …☆12Jul 27, 2026Updated 2 months ago
- Speech-to-text transcription VST3/ARA plugin☆63Jun 8, 2026Updated 3 months ago
- Voice activity detection and speaker gender segmentation audiovisual corpus☆16Jan 20, 2025Updated last year
- Human body part segmentation model, trained with 22 class labels.☆17Sep 28, 2023Updated 2 years ago
- Single Image Haze Removal Using AODNet in Pytorch☆15Mar 5, 2021Updated 5 years ago
- A Java port of whisper 3, based on the huggingface version, using DJL.☆19Apr 3, 2024Updated 2 years ago
- ☆10Dec 10, 2021Updated 4 years ago
- Comparing Audio Features for Unsupervised Sound Classification☆10Jun 22, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 这是一个简单的工具,用于方便地将Typora编辑器中的图片上传至Alist云存储服务☆12Oct 8, 2025Updated 11 months ago
- Whisfusion: Parallel ASR Decoding via a Diffusion Transformer☆31Aug 22, 2025Updated last year
- Trainer and Evaluation scripts for fine-tuning Whisper models for the Ukrainian language☆23Jan 13, 2023Updated 3 years ago
- Transform audio files into mel spectrograms for text-to-speech model training☆12Aug 25, 2021Updated 5 years ago
- ☆12May 1, 2019Updated 7 years ago
- Official source code for the paper "Tailored Design of Audio-Visual Speech Recognition Models using Branchformers"☆15Feb 24, 2025Updated last year
- WaveGANによる音声生成器☆13Feb 9, 2024Updated 2 years ago
- A simple RSS Feed Reader based on web technologies (HTML, CSS, JavaScript)☆13Jan 30, 2026Updated 7 months ago
- 一个强调工程化、可观测、可测试、可扩展 的 RAG 项目。TraceRAG 的目标不是只把答案“生成出来”,而是把文档导入、切块、向量化、检索、带来源回答、评估与后续 tracing 拆成可独立验证的阶段,逐步演进成一个可维护、可解释、可复盘的生产级 RAG。☆15Apr 2, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆24Jul 17, 2020Updated 6 years ago
- ☆13Aug 25, 2021Updated 5 years ago
- SGen is a generator capable of producing efficient hardware designs operating on streaming datasets. “Streaming” means that the dataset i…☆28Nov 11, 2025Updated 10 months ago
- 中国法律快查手册☆12Aug 19, 2025Updated last year
- Code for InterSpeech 2024 Paper: LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition☆19Jul 16, 2024Updated 2 years ago
- Resnet50 Quantization for Inference Speedup in PyTorch☆23Jan 30, 2021Updated 5 years ago
- Collaborative audio annotation tool☆17Sep 16, 2022Updated 4 years ago
- VHDL package for reading formatted data from comma-separated-values (CSV) files☆23Sep 10, 2013Updated 13 years ago
- Encode an image to sound (WAV file) and view it as a spectrogram. Optimized Python 3 version.☆11Jan 25, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Prepare spectrograms from audio for training a Riffusion model☆16Mar 6, 2023Updated 3 years ago
- A collection of helper scripts for Clojure, Java, Ledger and Taskwarrior. Written in Clojure.☆13Jun 2, 2023Updated 3 years ago
- Rainbowgram with Python☆13Jan 28, 2019Updated 7 years ago
- ☆19May 9, 2019Updated 7 years ago
- Keras implementation of conditional waveGAN. Application to knocking sound effects with emotion.☆10Jun 22, 2020Updated 6 years ago
- Grad-CAM (Gradient-weighted Class Activation Mapping)☆13Dec 20, 2019Updated 6 years ago
- Basic wavenet and fftnet vocoder model.☆19Feb 7, 2022Updated 4 years ago