[ICML26] AVGen-Bench is a task-driven benchmark for multi-granular evaluation of Text-to-Audio-Video (T2AV) generation.
☆29Jul 2, 2026Updated last month
Alternatives and similar repositories for AVGen-Bench
Users that are interested in AVGen-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆22Apr 24, 2026Updated 3 months ago
- CastleHill: Separable Causal Diffusion / Varitaion Flow Maps for LTX-2 long-form video generation☆15May 19, 2026Updated 3 months ago
- A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation☆34Jun 9, 2026Updated 2 months ago
- ☆17Jul 8, 2026Updated last month
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆47Jul 26, 2026Updated 3 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- NVV-SuperBench: Beyond Words, Beyond Quality—Benchmarking Nonverbal Vocalizations in Speech Generation (Interspeech 2026 long paper)☆18Jun 21, 2026Updated 2 months ago
- OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video☆19Apr 24, 2026Updated 4 months ago
- [IEEE TCSVT] Preprocessing Enhanced Image Compression for Machine Vision☆19Mar 23, 2025Updated last year
- Official repo for Directional Self-supervised Learning for Heavy Image Augmentations [CVPR2022]☆12Jun 29, 2022Updated 4 years ago
- [ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.☆381Mar 29, 2026Updated 4 months ago
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Aug 4, 2026Updated 2 weeks ago
- [ECCV 2026 Oral] Official implementation of "OmniForcing: Unleashing Real-time Joint Audio-Visual Generation"[arXiv:2603.11647]. OmniForc…☆182Jul 23, 2026Updated last month
- ☆11Mar 4, 2025Updated last year
- Code for "OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation"☆153Jun 18, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- HippoMM: Hippocampal-inspired Multimodal Memory☆22May 22, 2025Updated last year
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models☆33Feb 3, 2026Updated 6 months ago
- Official Code of NAVA: Native Audio-Visual Alignment for Generation.☆226Jun 30, 2026Updated last month
- ASLP Summer Inter@NPU☆13Jul 30, 2024Updated 2 years ago
- ☆19Jun 13, 2026Updated 2 months ago
- [NeurIPS 2024] The official implementation of "Image Copy Detection for Diffusion Models"☆18Oct 1, 2024Updated last year
- [ECCV 2024] RGBD GS-ICP SLAM☆14Nov 5, 2024Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- This respository is used for time reasoning task for mult-session dialogue system.☆18Feb 7, 2026Updated 6 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official respository for ReasonGen-R1☆74Jun 23, 2025Updated last year
- MTVCraft: An Open Veo3-style Audio-Video Generation Demo☆99Oct 8, 2025Updated 10 months ago
- 🏠 [ECCV 2024] The core gsplat component for GaussianImage☆42Mar 13, 2026Updated 5 months ago
- [ICML'26] Toward Human-like Audio-Visual Intelligence of Omni-MLLMs☆18Jun 20, 2026Updated 2 months ago
- Dashboard for orchestrating many concurrent Claude Code / Codex windows — triage, full-text search, skill/memory analytics☆49Jul 27, 2026Updated 3 weeks ago
- [CVPR 2025] T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation☆123Oct 25, 2025Updated 9 months ago
- Awesome GAN-based Image Restoration☆12Mar 11, 2024Updated 2 years ago
- ☆23Mar 2, 2026Updated 5 months ago
- [ICLR26] Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs☆33Dec 9, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LEMAS‑Edit is a multilingual speech editing system, supporting 10 languages: Chinese English Spanish Russian French German Italian Portug…☆21Mar 31, 2026Updated 4 months ago
- ☆59Apr 22, 2026Updated 4 months ago
- [NeurIPS 2025 Spotlight] Demo implementation of MoCha Towards Movie-Grade Talking Character Synthesis☆17Dec 27, 2025Updated 7 months ago
- ☆88Nov 16, 2025Updated 9 months ago
- ☆27Sep 10, 2025Updated 11 months ago
- [ICCV'25] Official PyTorch Implementation of "VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models"☆17Dec 8, 2025Updated 8 months ago
- ☆11Mar 11, 2025Updated last year