Official implementation of "Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound". IEEE TASLP 2025.
☆19Feb 27, 2026Updated 7 months ago
Alternatives and similar repositories for video-foley
Users that are interested in video-foley are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation of V-AURA: Temporally Aligned Audio for Video with Autoregression (ICASSP 2025) (Oral)☆36Feb 11, 2026Updated 7 months ago
- Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos☆26Oct 1, 2024Updated last year
- Official PyTorch implementation of "Conditional Generation of Audio from Video via Foley Analogies".☆93Dec 8, 2023Updated 2 years ago
- Stable-V2A: Synthesis of Synchronized Sound Effect with Temporal and Semantic Controls☆18May 27, 2025Updated last year
- LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation (INTERSPEECH 2024)☆44Jun 13, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆45Jan 13, 2025Updated last year
- Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation☆32Mar 8, 2024Updated 2 years ago
- The official implementation of the IJCAI 2024 paper "MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models".☆49Sep 11, 2024Updated 2 years ago
- The official repo for Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation☆66Jul 2, 2025Updated last year
- Implementation of "Factorized Diffusion Autoencoder for Unsupervised Disentangled Representation Learning" (AAAI 2024)☆17Dec 18, 2023Updated 2 years ago
- [CVPR 2025] Repository of VidMuse☆150Jun 7, 2025Updated last year
- A solution to denoising and separating for two-speaker-mixed noisy speech, using a BSRNN inspired network.☆15Aug 22, 2023Updated 3 years ago
- [NeurIPS'25] Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders☆16May 28, 2025Updated last year
- This repository provides an easy way to train your models on the datasets of DCASE task 1.☆20May 28, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- POMA-3D: The Point Map Way to 3D Scene Understanding.☆16Nov 9, 2025Updated 10 months ago
- ☆37Jan 6, 2026Updated 8 months ago
- Official PyTorch implementation of ReWaS (AAAI'25) "Read, Watch and Scream! Sound Generation from Text and Video"☆45Dec 13, 2024Updated last year
- to release the source code for reproducing the results reported in our paper: https://arxiv.org/abs/2409.17550☆14Nov 15, 2024Updated last year
- [ICCV 2025] This repo is the official implementation of "Music Grounding by Short Video"☆26Sep 9, 2025Updated last year
- Submission for task 2 "First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring" of the DCASE challenge 2023 (h…☆18May 22, 2023Updated 3 years ago
- [BMVC 2022] Information Theoretic Representation Distillation☆19Oct 6, 2023Updated 2 years ago
- Ego4DSounds: A diverse egocentric dataset with high action-audio correspondence☆21Jun 14, 2024Updated 2 years ago
- Official repository supporting the L3DAS23 IEEE ICASSP Grand Challenge☆18Feb 10, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ASR text preprocessing utility☆21Aug 5, 2024Updated 2 years ago
- ☆82Mar 14, 2025Updated last year
- Official implementation of the pipeline presented in I hear your true colors: Image Guided Audio Generation☆125Jan 18, 2023Updated 3 years ago
- Just another FastSpeech 2 but cleaner code :)☆29Jun 28, 2024Updated 2 years ago
- [AAAI 2024] Understanding the Role of the Projector in Knowledge Distillation☆20Feb 13, 2024Updated 2 years ago
- ☆17Jun 10, 2025Updated last year
- Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models☆206May 29, 2024Updated 2 years ago
- The implementation for "Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions"☆51Apr 7, 2025Updated last year
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols☆20Nov 19, 2025Updated 10 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- 3D Gaussian Splat Easily Attacked to Cause Harm☆13Aug 5, 2025Updated last year
- Pytorch Implementation of the Model from "MIRASOL3B: A MULTIMODAL AUTOREGRESSIVE MODEL FOR TIME-ALIGNED AND CONTEXTUAL MODALITIES"☆26Jan 27, 2025Updated last year
- collection with description of super-resolution related papers, repositories, datasets, loss functions and etc.☆11Dec 12, 2023Updated 2 years ago
- A python package to extract information from MIDI files☆14Aug 30, 2023Updated 3 years ago
- 🔥 [ICLR 2025] Official PyTorch Model "Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark"☆26Feb 9, 2025Updated last year
- ☆26Aug 14, 2025Updated last year
- A standardized toolkit of Kernel Audio Distance (KAD)—a distribution-free, unbiased, and computationally efficient metric for evaluating …☆106Updated this week