Transformer-based online speech recognition system with TensorFlow 2
☆26Jan 22, 2021Updated 5 years ago
Alternatives and similar repositories for Taris
Users that are interested in Taris are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Audio-Visual Speech Recognition using Sequence to Sequence Models☆84Jul 10, 2020Updated 6 years ago
- Google Summer of Code 2017 Project: Development of Speech Recognition Module for Red Hen Lab☆44Aug 29, 2017Updated 8 years ago
- A PyTorch implementation of the Deep Audio-Visual Speech Recognition paper.☆244Feb 15, 2024Updated 2 years ago
- Python toolkit for Visual Speech Recognition☆37Jun 10, 2020Updated 6 years ago
- Audio Visual Speech Recognition☆23Aug 9, 2017Updated 9 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Dual cross modality attention audio-visual speech recognition model based on vgg transformer with hybrid CTC/attention architecture using…☆15Jul 2, 2020Updated 6 years ago
- PyTorch implementation of automatic speech recognition models.☆38Jan 10, 2021Updated 5 years ago
- Pytorch code for End-to-End Audiovisual Speech Recognition☆182Nov 18, 2022Updated 3 years ago
- Voice Music Separation competing for 6th Huawei Cup in ZJU☆11Jun 2, 2015Updated 11 years ago
- Scripts for computing common lyrics-to-audio alignment evaluation metrics. Usable evaluation for any token-based alignment (e.g. if tok…☆18Oct 27, 2020Updated 5 years ago
- Audio-Visual Speech Recognition using Deep Learning☆61Nov 14, 2018Updated 7 years ago
- A tool to collect/validate audio recordings from workers on Amazon Mechanical Turk. Written in Python/Flask. (originally hosted on github…☆16Dec 19, 2022Updated 3 years ago
- A SPMI Lab toolkit for language models.☆11Apr 12, 2017Updated 9 years ago
- MTGAN: Speaker Verification through Multitasking Triplet Generative Adversarial Networks☆19Feb 29, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 56 language, 1 model Multilingual ASR☆25Jul 25, 2021Updated 5 years ago
- Cached Multi-Lora Composition for Multi-Concept Image Generation☆17Jun 13, 2025Updated last year
- This is an extension of kaldi speech recognition software which allows to perform decoding of speech with hybrid word and phoneme graphs.…☆11Feb 4, 2020Updated 6 years ago
- PyTorch implementation of "Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video" (ICCV2021)☆22Apr 11, 2022Updated 4 years ago
- tf 2.0 implementation of Listen, attend and spell☆21Jan 19, 2021Updated 5 years ago
- ☆37Dec 23, 2020Updated 5 years ago
- ☆10Jun 2, 2021Updated 5 years ago
- This is the official code for the paper "Reconstruct before Query: Continual Missing Modality Learning with Decomposed Prompt Collaborati…☆12Aug 13, 2024Updated 2 years ago
- ICASSP'22 Training Strategies for Improved Lip-Reading; ICASSP'21 Towards Practical Lipreading with Distilled and Efficient Models; ICASS…☆438May 18, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Tensorflow 2 Speech Recognition Code (Transformer)☆25Jun 29, 2020Updated 6 years ago
- Code for the paper: Audio-Visual Model Distillation Using Acoustic Images☆21Mar 24, 2023Updated 3 years ago
- Implementation of meta-transfer-learning for ASR and LM (ACL 2020)☆52Jul 30, 2020Updated 6 years ago
- The speaker-labeled information of LRW dataset, which is the outcome of the paper "Speaker-adaptive Lip Reading with User-dependent Paddi…☆10Oct 12, 2023Updated 2 years ago
- Script for converting kaldi GMM/HMM models to HTK format☆11Jul 18, 2024Updated 2 years ago
- Code repository for the Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR) dataset.☆42Jul 16, 2024Updated 2 years ago
- ChiNese Text Normalization (CNTN) tool for Text-to-speech system☆37Apr 12, 2018Updated 8 years ago
- An example directory for running Multi-Task Learning training on Kaldi neural networks. In Kaldi-speak, this is an egs dir for nnet3 trai…☆55Jan 2, 2020Updated 6 years ago
- ☆12Apr 16, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official git for "Fast Affine Motion Estimation for Versatile Video Coding (VVC) Encoding"☆11Sep 14, 2020Updated 5 years ago
- Project to segment video stream into separate shots☆13Oct 30, 2018Updated 7 years ago
- Multistream CNN for Robust Acoustic Modeling☆40Jun 17, 2021Updated 5 years ago
- ☆277Jan 15, 2021Updated 5 years ago
- streaming attention networks for end-to-end automatic speech recognition☆56May 6, 2020Updated 6 years ago
- Transformer implementation speciaized in speech recognition tasks using Pytorch.☆65Nov 28, 2021Updated 4 years ago
- An upgrade framework for train and validate compare with icefall using Lightning.☆16Mar 26, 2025Updated last year