[CVPR25] Official Implementation of CAV-MAE Sync
☆32Apr 5, 2026Updated 4 months ago
Alternatives and similar repositories for cav-mae-sync
Users that are interested in cav-mae-sync are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VGGSounder, a multi-label audio-visual classification dataset with modality annotations.☆17Jun 30, 2026Updated 2 months ago
- Implementation of the paper: "Audio Mamba: Bidirectional State Space Model for Audio Representation Learning" in pytorch☆15Updated this week
- ☆12Mar 12, 2023Updated 3 years ago
- ☆82Mar 14, 2025Updated last year
- VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.☆18May 31, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Solos: A Dataset for Audio-Visual Music Analysis☆24Feb 17, 2023Updated 3 years ago
- Unsupervised word segmentation and clustering of speech☆13Feb 17, 2017Updated 9 years ago
- [ICLR2025] Frechet Wavelet Distance: A metric to detect domain bias in Generative models.☆18Sep 2, 2025Updated 11 months ago
- Temperature Schedules for self-supervised contrastive methods on long-tail data (ICLR'23)☆18Apr 25, 2023Updated 3 years ago
- Unified Multisensory Perception: Weakly-Supervised Audio-Visual Video Parsing, ECCV, 2020. (Spotlight)☆90Jul 25, 2024Updated 2 years ago
- ☆23Dec 5, 2023Updated 2 years ago
- [CHIL 2024] Interpretation of Intracardiac Electrograms Through Textual Representations☆12Sep 4, 2024Updated last year
- ☆15Mar 27, 2025Updated last year
- Segmentation of prostatic zones (peripheral zone, central gland, AFS and distal prostatic urethra )☆19Apr 12, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation (ACM MM 2024)☆20Mar 17, 2025Updated last year
- Uncertainty-Guided Pseudo-Labelling with Model Averaging☆11Mar 17, 2026Updated 5 months ago
- [CVPR 2025] Pytorch implementation of the paper "Learning to Highlight Audio by Watching Movies"☆15Oct 1, 2025Updated 10 months ago
- ☆12Mar 24, 2024Updated 2 years ago
- awesome-audio-visual-robustness☆11Jan 27, 2024Updated 2 years ago
- ☆36Jan 20, 2025Updated last year
- Code for the AVLnet (Interspeech 2021) and Cascaded Multilingual (Interspeech 2021) papers.☆54Mar 30, 2022Updated 4 years ago
- Video descriptions of research papers relating to foundation models and scaling☆29Mar 16, 2023Updated 3 years ago
- Visual Speech Recognition For Low-Resource Languages with Automatic Labels (ICASSP 2024)☆17Mar 17, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICMR 2025] Official Repository for The Paper, Let Network Decide What to Learn: Symbolic Music Understanding Model Based on Large-scale …☆19Aug 17, 2025Updated last year
- Offical code for the CVPR 2024 Paper: Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language☆90Jun 12, 2024Updated 2 years ago
- WildVSR☆22Dec 13, 2023Updated 2 years ago
- Official code for the paper "Scaling Multilingual Visual Speech Recognition"☆20Aug 15, 2025Updated last year
- Official implementation of "VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification"☆26May 7, 2026Updated 3 months ago
- Official repository for the MMFM challenge☆26Jun 18, 2024Updated 2 years ago
- ☆23Jun 17, 2025Updated last year
- Pitman-Yor processes in python☆26Apr 18, 2014Updated 12 years ago
- 哈工大2021秋计算机网络☆13Mar 30, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- First neural GPT aligned with text and speech. Welcome to join us to make better foundation model in neural modality.☆14Oct 30, 2024Updated last year
- ICCV 2021☆34May 11, 2022Updated 4 years ago
- Download audioset data super fastly with youtube-dl, ffmpeg and python multiprocessing☆48Aug 1, 2024Updated 2 years ago
- PyTorch DataSet and Jupyter demos for MusicNet☆75Mar 21, 2024Updated 2 years ago
- ☆12Feb 27, 2024Updated 2 years ago
- Continual Online Recalibration with Pseudo-labels☆15Jun 20, 2024Updated 2 years ago
- A reconstruction framework for materializing subjective experiences from brain signals☆15Jan 18, 2025Updated last year