[CVPR25] Official Implementation of CAV-MAE Sync
☆31Apr 5, 2026Updated 3 months ago
Alternatives and similar repositories for cav-mae-sync
Users that are interested in cav-mae-sync are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VGGSounder, a multi-label audio-visual classification dataset with modality annotations.☆17Jun 30, 2026Updated 3 weeks ago
- Implementation of the paper: "Audio Mamba: Bidirectional State Space Model for Audio Representation Learning" in pytorch☆15Updated this week
- ☆12Mar 12, 2023Updated 3 years ago
- VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.☆18May 31, 2026Updated last month
- Code for the C2KD paper (ICASSP 2023)☆20May 15, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Unsupervised word segmentation and clustering of speech☆13Feb 17, 2017Updated 9 years ago
- [ICLR2025] Frechet Wavelet Distance: A metric to detect domain bias in Generative models.☆18Sep 2, 2025Updated 10 months ago
- Temperature Schedules for self-supervised contrastive methods on long-tail data (ICLR'23)☆18Apr 25, 2023Updated 3 years ago
- Unified Multisensory Perception: Weakly-Supervised Audio-Visual Video Parsing, ECCV, 2020. (Spotlight)☆90Jul 25, 2024Updated last year
- [CHIL 2024] Interpretation of Intracardiac Electrograms Through Textual Representations☆12Sep 4, 2024Updated last year
- [ICCV25] Official Implementation of LeGrad☆98Oct 14, 2024Updated last year
- ☆15Mar 27, 2025Updated last year
- Fast training of unitary deep network layers from low-rank updates☆30Dec 11, 2022Updated 3 years ago
- This is an official pytorch implementation of Learning To Recognize Procedural Activities with Distant Supervision. In this repository, w…☆43Feb 21, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 👁️🗨️ Scientists often do the same bad stuff. Automate giving deterministic feedback during peer review with determinstic (LLM-free)☆32May 9, 2025Updated last year
- Segmentation of prostatic zones (peripheral zone, central gland, AFS and distal prostatic urethra )☆19Apr 12, 2021Updated 5 years ago
- Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation (ACM MM 2024)☆20Mar 17, 2025Updated last year
- Uncertainty-Guided Pseudo-Labelling with Model Averaging☆11Mar 17, 2026Updated 4 months ago
- 👆PyTorch Implementation of JEDi Metric described in "Beyond FVD: Enhanced Evaluation Metrics for Video Generation Quality"☆35Dec 8, 2024Updated last year
- [CVPR 2025] Pytorch implementation of the paper "Learning to Highlight Audio by Watching Movies"☆15Oct 1, 2025Updated 9 months ago
- ☆12Mar 24, 2024Updated 2 years ago
- Data Release for VALUE Benchmark☆30Feb 16, 2022Updated 4 years ago
- awesome-audio-visual-robustness☆11Jan 27, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for “ACE-HGNN: Adaptive Curvature ExplorationHyperbolic Graph Neural Network”☆17Mar 3, 2022Updated 4 years ago
- Code for the AVLnet (Interspeech 2021) and Cascaded Multilingual (Interspeech 2021) papers.☆54Mar 30, 2022Updated 4 years ago
- Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".☆292Mar 20, 2024Updated 2 years ago
- Video descriptions of research papers relating to foundation models and scaling☆29Mar 16, 2023Updated 3 years ago
- Visual Speech Recognition For Low-Resource Languages with Automatic Labels (ICASSP 2024)☆17Mar 17, 2025Updated last year
- [NAACL'25] Contains code and documentation for our VANE-Bench paper.☆24Aug 19, 2025Updated 11 months ago
- Repository for the paper: dense and aligned captions (dac) promote compositional reasoning in vl models☆28Nov 29, 2023Updated 2 years ago
- Humans-in-Kitchens Dataset API ([NeurIPS 2023 Dataset and Benchmark Track])☆41Sep 29, 2024Updated last year
- WildVSR☆22Dec 13, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official code for the paper "Scaling Multilingual Visual Speech Recognition"☆20Aug 15, 2025Updated 11 months ago
- Official repository for the MMFM challenge☆26Jun 18, 2024Updated 2 years ago
- 哈工大2021秋计算机网络☆13Mar 30, 2023Updated 3 years ago
- Chinese-Mimi 是对 Moshi 模型的声码器进行了中文语料上的适配。☆36Mar 13, 2025Updated last year
- DO with Terraform and Ansible☆11Jun 5, 2018Updated 8 years ago
- First neural GPT aligned with text and speech. Welcome to join us to make better foundation model in neural modality.☆14Oct 30, 2024Updated last year
- ICCV 2021☆34May 11, 2022Updated 4 years ago