In this work I investigate the speech command task developing and analyzing deep learning models. The state of the art technology uses convolutional neural networks (CNN) because of their intrinsic nature of learning correlated represen- tations as is the speech. In particular I develop different CNNs trained on the Google Speech Command Dataset…
☆20Jul 18, 2018Updated 8 years ago
Alternatives and similar repositories for learning_invariances_in_speech_recognition
Users that are interested in learning_invariances_in_speech_recognition are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Text-to-Speech Synthesis by Generating Spectrograms using Generative Adversarial Network☆10Dec 12, 2018Updated 7 years ago
- This repo contains script to download MUSIC dataset from youtube☆13Jan 19, 2024Updated 2 years ago
- Google Speech Command Dataset Classification Neural Network, CNN, RNN☆26Aug 29, 2017Updated 9 years ago
- ☆10Mar 22, 2022Updated 4 years ago
- Privacy-preserving Voice Analysis via Disentangled Representations☆13Aug 30, 2021Updated 5 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- TensorFlowLiteNet allows to use TensorFlowLite from C#.☆11Apr 14, 2021Updated 5 years ago
- Implemented 3 neural network architectures: 1) Combination of RNN LSTM nodes and CNN, 2) CNN with residual blocks similar to ResNet, 3) D…☆25Jan 19, 2018Updated 8 years ago
- Automatic Arabic diacritics restoration tool.☆19Aug 12, 2021Updated 5 years ago
- Following research on S4 in jax☆16Jun 15, 2022Updated 4 years ago
- Sinatra app for serving web fonts easily with proper caching and access-control headers☆15Sep 5, 2011Updated 15 years ago
- FEERCI: A Package for Fast non-parametric confidence intervals for Equal Error Rates☆12Mar 13, 2024Updated 2 years ago
- ☆13Aug 21, 2026Updated last month
- there are UKIJ and Uighursoft fonts☆13Oct 21, 2022Updated 3 years ago
- ☆11Jun 15, 2022Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Configuration files for my setup☆10Jun 4, 2023Updated 3 years ago
- uyghur text resource crawled from website☆12Dec 25, 2015Updated 10 years ago
- Incorporating AutoVocoder to MB-iSTFT-VITS☆47Dec 1, 2022Updated 3 years ago
- Target speaker automatic speech recognition (TS-ASR)☆15Oct 14, 2023Updated 2 years ago
- ☆33Nov 27, 2021Updated 4 years ago
- A query by humming system based on locality sensitive hashing indexes☆12May 8, 2014Updated 12 years ago
- Transformer based ASR Engine.☆13Aug 23, 2021Updated 5 years ago
- Time delay neural network (TDNN) implementation in Pytorch using unfold method☆207Nov 21, 2019Updated 6 years ago
- Maintained at https://github.com/crystal-lang-tools/emacs-crystal-mode☆17Nov 7, 2017Updated 8 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- dashboard module for emoncms☆14Sep 17, 2026Updated last week
- NNSVS向けの教師データのラベル作成支援ツールです。☆10Apr 5, 2023Updated 3 years ago
- Volcengine TOS C++ SDK☆14Sep 15, 2026Updated last week
- A parser and serializer to make it easier to manipulate .vtt files.☆12Sep 23, 2014Updated 12 years ago
- A punctuation transcription model to automatically add punctuation marks in an unpunctuated sentence or sentences.☆15Aug 6, 2020Updated 6 years ago
- Make N-Gram for Uyghur language☆15Dec 24, 2020Updated 5 years ago
- ☆29Sep 9, 2026Updated 2 weeks ago
- OBD Scan Tool .NET 2.0☆14Oct 5, 2015Updated 10 years ago
- Train neural network via pytorch, and run nn model on ESP32☆11Dec 1, 2022Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Music Catalogizer + MP3 ID tag parser + Radio (WPF, WebApi, Angular)☆14Oct 20, 2021Updated 4 years ago
- C# source code for creating MotionJPEG.☆16Jan 9, 2020Updated 6 years ago
- ☆11Nov 13, 2015Updated 10 years ago
- This is a single-speaker neural text-to-speech (TTS) system capable of training in a end-to-end fashion. It is inspired by the Tacotron a…☆12Dec 28, 2018Updated 7 years ago
- Ruby wrapper for the arXiv API☆27May 13, 2026Updated 4 months ago
- Master thesis of Ondrej Platek: Automatic speech recognition using Kaldi. Supervised by Filip Jurcicek.☆15Feb 20, 2020Updated 6 years ago
- ☆13Oct 12, 2018Updated 7 years ago