Multimodal Speech Recognition for phoneme level prediction using Audio-Visual data from TCDTIMIT dataset implementing RNNs with LSTMs for the audio subnetwork and CNN-LSTMs for the video subnetwork.
☆15Jul 27, 2023Updated 3 years ago
Alternatives and similar repositories for Lipreading-Using-Mutimodal-Speech-Recognition
Users that are interested in Lipreading-Using-Mutimodal-Speech-Recognition are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Aug 8, 2023Updated 3 years ago
- Dual cross modality attention audio-visual speech recognition model based on vgg transformer with hybrid CTC/attention architecture using…☆15Jul 2, 2020Updated 6 years ago
- processing and extracting of face and mouth image files out of the TCDTIMIT database☆47Sep 22, 2020Updated 5 years ago
- Python toolkit for Visual Speech Recognition☆37Jun 10, 2020Updated 6 years ago
- End to End Multiview Lip Reading☆10Jan 26, 2018Updated 8 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆24May 11, 2025Updated last year
- Aty-TTS: Improving fairness for spoken language understanding in atypical speech with Text-to-Speech☆12May 14, 2025Updated last year
- 🎮 Use a Raspberry Pi to control a LoPy over UART☆12Mar 9, 2017Updated 9 years ago
- Examples of how to use a scriptable object to setup blendshape mapping for ARKit Blendshapes☆15Feb 17, 2021Updated 5 years ago
- ☆17Aug 9, 2023Updated 3 years ago