AutoCorpus is a set of utilities that enable automatic extraction of language corpora and language models from publicly available datasets. Autocorpus utilities follow the Unix design philosophy and integrate easily into custom data processing pipelines.
☆37Feb 1, 2012Updated 14 years ago
Alternatives and similar repositories for AutoCorpus
Users that are interested in AutoCorpus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This will hold the crowdsourcing platform to be used to store voice data from various speakers which will act as input dataset for speech…☆17Mar 6, 2023Updated 3 years ago
- This is application for dysarthria to improve their pronunciation by using deep learning☆10Dec 29, 2020Updated 5 years ago
- Perform the forced decoding with target transcription☆11Sep 12, 2018Updated 7 years ago
- steps to perform text-based speaker diarization with kaldi toolkit☆12Nov 2, 2018Updated 7 years ago
- Speech Processing & Linguistic Analysis Tool☆11Jun 30, 2019Updated 7 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Coqui STT (🐸STT) based forced alignment tool☆13Feb 24, 2022Updated 4 years ago
- ☆13Jun 30, 2026Updated 2 months ago
- ☆18Apr 28, 2021Updated 5 years ago
- This is a mirror of https://gitlab.com/tiro-is/tiro-speech-core☆15Jun 19, 2023Updated 3 years ago
- Grapheme to phoneme toolkit using joint-modelling + CRFs in java☆16Jul 14, 2018Updated 8 years ago
- A set of scripts to use in preparing a corpus for speech-to-text processing with the Kaldi Automatic Speech Recognition Library.☆15May 19, 2020Updated 6 years ago
- Tools for working with the CMU Pronunciation Dictionary☆36Sep 5, 2017Updated 8 years ago
- finite-state toolkit, EM and Bayesian (Gibbs sampling) training for FST and context-free derivation forests☆15Jan 24, 2017Updated 9 years ago
- wake word spotting with kaldi☆19Dec 3, 2020Updated 5 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Docker image and scripts for training finetuned or completely personal Kaldi speech models. Particularly for use with kaldi-active-gramma…☆21Jan 24, 2022Updated 4 years ago
- Deploy Kaldi models using grpc for bidirectional streaming.☆17Sep 30, 2024Updated last year
- An app that graphs and compares the pitch contours of spoken language, to help language learners perfect their intonation (Hackbright Spr…☆32Jul 20, 2017Updated 9 years ago
- This app is intended to automatically create a corpus for ASR systems using pseudo-labeling.☆27Feb 15, 2024Updated 2 years ago
- BurrMill core☆22Nov 2, 2021Updated 4 years ago
- ☆22Jul 8, 2021Updated 5 years ago
- Easier analysis of large speech corpora☆25Jun 22, 2021Updated 5 years ago
- Application boilerplates for Restler. Each branch contains a flavor, find the one that suits you.☆11Jun 13, 2021Updated 5 years ago
- Python based & web based IDE for DTrace with Data Visualizations☆15Jun 13, 2012Updated 14 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- 📖 LanMIT: A Toolkit for Improving Language Models in Low-resourced Speech Recognition based on Kaldi.☆22Jul 12, 2019Updated 7 years ago
- A free & open tool for transcribing audio interviews with offline ASR support☆25Dec 21, 2023Updated 2 years ago
- A handy dataset of noises for ASR☆22May 29, 2019Updated 7 years ago
- Discussion Summarization is the process of condensing a text document which is a collection of discussion threads, using CBS (Cluster Bas…☆12Apr 10, 2014Updated 12 years ago
- Phonetic and phonological vocoding platform☆17Nov 23, 2016Updated 9 years ago
- ☆33Nov 27, 2021Updated 4 years ago
- Sparkline control for WPF and Silverlight☆12Aug 5, 2011Updated 15 years ago
- Java interfaces and tools for Kaldi speech recognition.☆22Oct 2, 2016Updated 9 years ago
- autohotkey syntax highlighting☆14Oct 19, 2010Updated 15 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code from the paper Reflection for the Masses by Charlotte Herzeel, Pascal Costanza, and Theo D'Hondt.☆15Jun 21, 2021Updated 5 years ago
- ☆25Jun 14, 2022Updated 4 years ago
- Implementation of the MMDAgent for use as a live receptionist in Carnegie Mellon's School of Computer Science.☆16Apr 11, 2013Updated 13 years ago
- Implementation of different noise embeddings for noise aware training of Kaldi acoustic models.☆13Feb 13, 2021Updated 5 years ago
- A visualizer for multi-dimensional semantic data☆38Oct 24, 2011Updated 14 years ago
- small python app to help practice speech shadowing, helpful for language learning☆16Jun 25, 2020Updated 6 years ago
- Some utility commands you wish were in Python's pip☆15Apr 2, 2014Updated 12 years ago