[ACM MM 2026] ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
☆153Aug 28, 2026Updated last month
Alternatives and similar repositories for controlfoley
Users that are interested in controlfoley are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆46Aug 20, 2026Updated last month
- 神棍☆17May 1, 2026Updated 5 months ago
- [Technical Report] An End-to-End Multimodal GUI Agent for Real Mobile Environments☆87Sep 18, 2026Updated 2 weeks ago
- end-to-end text to audio scene generation model☆53Jun 16, 2026Updated 3 months ago
- State-of-the-art continious audio tokenization☆42Mar 9, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICCV 2025] Implementation of the paper "Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs"☆83Oct 25, 2025Updated 11 months ago
- ☆29Apr 6, 2026Updated 5 months ago
- Generate a complete audio clip with music, intelligible speech, and sound effects from text in one pass.☆46May 27, 2026Updated 4 months ago
- Foley-Omni: a unified multimodal audio generation model for task-level synthesis and complete video soundtrack generation, producing spee…☆27Jun 5, 2026Updated 4 months ago
- ☆32Mar 27, 2026Updated 6 months ago
- ☆17Apr 30, 2026Updated 5 months ago
- Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation☆76Jun 16, 2026Updated 3 months ago
- Gemma-based Multilingual Machine Translation Models☆93Aug 14, 2026Updated last month
- Official code for "WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling"☆64Jun 27, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- interpretability work and exploration for krea☆60Jul 12, 2026Updated 2 months ago
- Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model☆29May 21, 2026Updated 4 months ago
- Official Implementation of GLAP - General Language Audio Pretraining☆76May 14, 2026Updated 4 months ago
- ☆44Jun 3, 2026Updated 4 months ago
- ☆52Apr 27, 2026Updated 5 months ago
- WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling☆212Jun 6, 2026Updated 3 months ago
- ☆23Jan 8, 2024Updated 2 years ago
- Ming-omni-tts: Simple and Efficient Unified Generation of Speech, Music, and Sound with Precise Control☆265Feb 26, 2026Updated 7 months ago
- ☆27May 26, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Keyword Spotting using BCResNet and Arcface Loss☆13Jan 28, 2022Updated 4 years ago
- ComfyUI批处理工具☆21Apr 6, 2026Updated 5 months ago
- Pushing the Frontier of Long Video Generation Standalone, inference-only release for minute-level multi-shot audio-video generation with…☆57Sep 16, 2026Updated 2 weeks ago
- [NeurIPS 2025] Implementation of the paper "BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent"☆19Nov 27, 2025Updated 10 months ago
- NVV-SuperBench: Beyond Words, Beyond Quality—Benchmarking Nonverbal Vocalizations in Speech Generation (Interspeech 2026 long paper)☆19Jun 21, 2026Updated 3 months ago
- A curated list of models, benchmarks, tools and guides for audio editing☆49Updated this week
- Audio skills for Claw☆29Apr 16, 2026Updated 5 months ago
- ☆454Mar 25, 2026Updated 6 months ago
- ComfyUI-AudioLoopHelper☆16Jul 5, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- X-Voice☆184Sep 24, 2026Updated last week