Code on selecting an action based on multimodal inputs. Here in this case inputs are voice and text.
☆73Jun 7, 2021Updated 5 years ago
Alternatives and similar repositories for Multimodal-action-recognition
Users that are interested in Multimodal-action-recognition are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Pytorch Implementation for Continual Learning For On-Device Environmental Sound Classification☆14Jul 19, 2022Updated 4 years ago
- Multimodal late fusion for deepfake detection using video and audio data☆11May 7, 2019Updated 7 years ago
- Pytorch implementation of DSR-RL for Video Summarization Task☆12Aug 30, 2021Updated 5 years ago
- Multimodal speech recognition using lipreading (with CNNs) and audio (using LSTMs). Sensor fusion is done with an attention network.☆69Nov 19, 2022Updated 3 years ago
- Chinese BERT classification with tf2.0 and audio classification with mfcc☆14Dec 2, 2020Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- collection of skeleton-based human action recognition☆10Jun 28, 2020Updated 6 years ago
- (2020) Video Classification Neural Network☆30Feb 18, 2020Updated 6 years ago
- A collection of multimodal datasets, and visual features for VQA and captionning in pytorch. Just run "pip install multimodal"☆85Feb 25, 2022Updated 4 years ago
- ☆14Aug 13, 2020Updated 6 years ago
- A search engine implementation using OpenAI's clip model☆10Jun 20, 2021Updated 5 years ago
- Pretraining summarization models using a corpus of nonsense☆13Sep 28, 2021Updated 4 years ago
- A Pytorch implementation of emotion recognition from videos☆18Sep 15, 2020Updated 6 years ago
- This repository contains various models targetting multimodal representation learning, multimodal fusion for downstream tasks such as mul…☆920Mar 15, 2023Updated 3 years ago
- ☆25Nov 23, 2021Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- PLLay: Efficient Topological Layer based on Persistence Landscapes☆23Dec 10, 2020Updated 5 years ago
- Predicting Political Instability and Social Conflicts Using Multimodal Data☆10Jun 6, 2016Updated 10 years ago
- 👻 Code and benchmark for our EMNLP 2023 paper - "FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions"☆63May 31, 2024Updated 2 years ago
- Code for paper "Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition"☆18Jun 21, 2023Updated 3 years ago
- PaddleSeq☆10Mar 28, 2023Updated 3 years ago
- A re-implementation of the CVPR19 paper Quantization Networks on CIFAR-10, MNIST and ImageNet☆10Aug 9, 2020Updated 6 years ago
- Zicx's Notebook.☆10Nov 7, 2025Updated 10 months ago
- acnn for text-independent speaker recognition☆10Feb 8, 2022Updated 4 years ago
- Attacks against proposed image encryption schemes☆10Apr 27, 2020Updated 6 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering (NeurIPS 2020)☆91Oct 24, 2022Updated 3 years ago
- ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning. In ICCV, 2021.☆64Nov 18, 2021Updated 4 years ago
- PROJETO FINAL - Comunicação por luz visível (VLC):Análise e síntese por simulação utilizando o MATLAB. Autores: Gabrielle Cristina de Sou…☆11Mar 1, 2021Updated 5 years ago
- Official pytorch implementation of I2I translation with low resolution conditioning☆23Sep 2, 2021Updated 5 years ago
- (Competition) 6th -- Scene-Text-Detection-and-Recognition.☆11Jun 14, 2022Updated 4 years ago
- FG2021: Cross Attentional AV Fusion for Dimensional Emotion Recognition☆34Nov 29, 2024Updated last year
- Codebase for the paper: "TIM: A Time Interval Machine for Audio-Visual Action Recognition"☆54Nov 7, 2024Updated last year
- Pytorch implementation of Meta-Learning for Short Utterance Speaker Recognition with Imbalance Length Pairs (Interspeech, 2020)☆73Sep 16, 2020Updated 6 years ago
- A game engine based on Direct3D 12, C++17, and .NET 9 for learning purposes.☆73Oct 7, 2025Updated 11 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- An official implementation for " UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation"☆365Jul 25, 2024Updated 2 years ago
- Source code for ScaleGrad☆19Dec 28, 2021Updated 4 years ago
- ☆19Jul 27, 2021Updated 5 years ago
- Video Transformer Network☆41Jun 8, 2021Updated 5 years ago
- PyTorch – SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models.☆62Jun 28, 2022Updated 4 years ago
- Simple phoenix setup for padded window management☆13Apr 25, 2018Updated 8 years ago
- Resources for: Cross-Lingual Disaster-related Multi-label Tweet Classification with Manifold Mixup (ACL SRW 2020)☆11Sep 9, 2021Updated 5 years ago