[NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”
☆38Jul 24, 2025Updated last year
Alternatives and similar repositories for video-to-audio-through-text
Users that are interested in video-to-audio-through-text are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR 2025] Pytorch implementation of the paper "Hearing Anywhere in Any Environment"☆34Sep 18, 2025Updated 10 months ago
- ☆31Feb 4, 2021Updated 5 years ago
- ☆12Apr 30, 2025Updated last year
- Repository for SAVVY(Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing) Benchmark and SAVVY model☆25May 30, 2026Updated 2 months ago
- Implementation of Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching (NeurIPS'24)☆63Apr 3, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆20Aug 11, 2025Updated 11 months ago
- code for "When Counterpoints Meet Chinese Folk Melody"☆11Feb 19, 2021Updated 5 years ago
- Open repository of simulated Room Impulse Responses (RIR) accompanying the paper "Hearing Anywhere in Any Environment"☆82Aug 11, 2025Updated 11 months ago
- ☆53Mar 24, 2026Updated 4 months ago
- The official implementation of V-AURA: Temporally Aligned Audio for Video with Autoregression (ICASSP 2025) (Oral)☆35Feb 11, 2026Updated 5 months ago
- A curated list of Vision (video/image) to Audio Generation☆107Feb 10, 2026Updated 5 months ago
- Explaining audio differences using language☆16Feb 11, 2025Updated last year
- [AAAI 2024] V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models☆29Dec 14, 2023Updated 2 years ago
- Official implementation of the CVPR 2026 paper "SonoWorld: From One Image to a 3D Audio-Visual Scene."☆41Jul 6, 2026Updated 3 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Repository for PREDICT & CLUSTER: Unsupervised Skeleton Based Action Recognition☆111Aug 28, 2023Updated 2 years ago
- ☆141Jan 24, 2026Updated 6 months ago
- small audio language model for reasoning☆88Dec 4, 2025Updated 8 months ago
- Official PyTorch implementation of ReWaS (AAAI'25) "Read, Watch and Scream! Sound Generation from Text and Video"☆44Dec 13, 2024Updated last year
- NeuroAI-UW seminar, a regular weekly seminar for the UW community, organized by NeuroAI Shlizerman Lab.☆58May 27, 2025Updated last year
- Solos: A Dataset for Audio-Visual Music Analysis☆24Feb 17, 2023Updated 3 years ago
- ☆16Jul 8, 2026Updated 3 weeks ago
- Official PyTorch implementation of "Conditional Generation of Audio from Video via Foley Analogies".☆93Dec 8, 2023Updated 2 years ago
- [CVPR 2024] Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners☆155Jul 6, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with syntheti…☆135Mar 3, 2026Updated 5 months ago
- [CVPR 2026] Official PyTorch implementation of SelVA "Hear What Matters! Text-conditioned Selective Video-to-Audio Generation"☆16Mar 27, 2026Updated 4 months ago
- Code for the paper "DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models"☆17Oct 23, 2025Updated 9 months ago
- ☆15Jul 20, 2026Updated 2 weeks ago
- [ICLR'25] MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation☆39Dec 25, 2025Updated 7 months ago
- Code for ICLR 2024 Paper: CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models☆23Jul 10, 2024Updated 2 years ago
- Official repository for the paper "Audio ControlNet for Fine-Grained Audio Generation and Editing".☆77Feb 7, 2026Updated 5 months ago
- [ICCV'25] Official PyTorch Implementation of "VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models"☆17Dec 8, 2025Updated 7 months ago
- [ICASSP2025] Official code for VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis☆52Apr 9, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Benchmarking for Audio-Text and Audio-Visual Generation; Supports FAD, FD_VGG, FD_PANNs, FD_PaSST, IS_PaSST, IS_PANNs, KL_PaSST, KL_PANNs…☆80Feb 14, 2026Updated 5 months ago
- ☆25Nov 25, 2025Updated 8 months ago
- ☆49Jul 10, 2024Updated 2 years ago
- GitHub repository for AudioToolAgent☆20Feb 13, 2026Updated 5 months ago
- Materials for "Multimedia Deepfake Detection" Tutorial @ ICME 2024☆17Aug 26, 2024Updated last year
- [ISMIR 2025] A curated list of vision-to-music generation: methods, datasets, evaluation and challenges.☆126Aug 9, 2025Updated 11 months ago
- Curated list for papers, codes and resources related to Text-to-Audio (TTA) Generation☆76Jul 20, 2026Updated 2 weeks ago