☆26Feb 10, 2026Updated 6 months ago
Alternatives and similar repositories for UniAudio2Demo
Users that are interested in UniAudio2Demo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A framework for camera-controllable image editing using unified geometric guidance and video models.☆65Jun 25, 2026Updated last month
- Sampling code for MobileWan☆97Updated this week
- [ICML2026] From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors☆93Apr 30, 2026Updated 3 months ago
- GitHub repository for AudioToolAgent☆20Feb 13, 2026Updated 5 months ago
- Code repo for EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation☆42Mar 6, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆15Oct 13, 2025Updated 9 months ago
- UniMesh: Unifying 3D Mesh Understanding and Generation☆57Jul 14, 2026Updated 3 weeks ago
- DreamStyle: A Unified Framework for Video Stylization☆124Jan 7, 2026Updated 7 months ago
- [CVPR 2026] When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models☆68Apr 11, 2026Updated 3 months ago
- Inference server for MioTTS, a lightweight and fast LLM-based TTS model.☆200Feb 14, 2026Updated 5 months ago
- Official PyTorch Implementation of Ctrl-Crash 💥☆53Jun 3, 2025Updated last year
- UniAudio 2.0: An audio fundation model for text, speech, sound, and music☆213Feb 14, 2026Updated 5 months ago
- Code for 'JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion'☆267May 11, 2026Updated 2 months ago
- [CVPR 2026] Official PyTorch implementation of SelVA "Hear What Matters! Text-conditioned Selective Video-to-Audio Generation"☆16Mar 27, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆18Nov 19, 2025Updated 8 months ago
- [ICCV 2025] Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping☆97Nov 30, 2025Updated 8 months ago
- [CVPR'26] VecGlypher: Unified Vector Glyph Generation with Language Models☆139Feb 26, 2026Updated 5 months ago
- A simple implementation for improving CosyVoice2 by GRPO method☆39May 5, 2026Updated 3 months ago
- [CVPR 2026] ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands☆132Apr 22, 2026Updated 3 months ago
- E2E TTS using Conditional Flow Matching (Experimental*)☆71Nov 10, 2023Updated 2 years ago
- Official repository of paper "ProEdit: Inversion-based Editing From Prompts Done Right"☆116Feb 5, 2026Updated 6 months ago
- ComfyUI custom nodes for Foundation-1 | Structured Text-to-Sample Diffusion for Music Production☆107Mar 22, 2026Updated 4 months ago
- Landing Page for Divide and Remaster v3☆26Jul 29, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- We propose a novel modular framework that learns to dynamically mix low-rank adapters (LoRAs) to improve visual analogy learning, enablin…☆75Aug 2, 2026Updated last week
- Learning to Simulate Mechanics in World Space☆56Jul 24, 2026Updated 2 weeks ago
- Adaptive Multimodal Reasoning via Reinforcement Learning☆23Jan 11, 2026Updated 6 months ago
- [CVPR 2026] Official repo of "MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing“☆110Apr 13, 2026Updated 3 months ago
- [CVPR 2026] PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow☆151Mar 28, 2026Updated 4 months ago
- [NeurIPS'25 Spotlight] Official implementation of "JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation"☆75Feb 26, 2026Updated 5 months ago
- [Official Repo] SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing☆215Apr 13, 2026Updated 3 months ago
- ☆15Jun 23, 2024Updated 2 years ago
- Official implementation of TISDiSS, a scalable framework for discriminative source separation.☆16Jul 31, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- SOTA Piano Transformer model trained on 4.2GB of Solo Piano MIDI music☆28Nov 9, 2023Updated 2 years ago
- Chorale Music Separation Dataset and Model Framework☆41Dec 5, 2022Updated 3 years ago
- [SIGGRAPH 2026 Journal] SegviGen: Repurposing 3D Generative Model for Part Segmentation☆161Mar 19, 2026Updated 4 months ago
- Code for the paper Proactive Hearing Assistants that Isolate Egocentric Conversations☆46Nov 19, 2025Updated 8 months ago
- ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation☆117Dec 11, 2025Updated 7 months ago
- ☆17Jun 13, 2025Updated last year
- ☆15Oct 27, 2025Updated 9 months ago