[ICML26] AVGen-Bench is a task-driven benchmark for multi-granular evaluation of Text-to-Audio-Video (T2AV) generation.
☆24Jul 2, 2026Updated last month
Alternatives and similar repositories for AVGen-Bench
Users that are interested in AVGen-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆21Apr 24, 2026Updated 3 months ago
- CastleHill: Separable Causal Diffusion / Varitaion Flow Maps for LTX-2 long-form video generation☆15May 19, 2026Updated 2 months ago
- A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation☆33Jun 9, 2026Updated last month
- ☆16Jul 8, 2026Updated 3 weeks ago
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆46Jul 26, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 🧂 [ECCV 2026] Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation☆16Apr 6, 2026Updated 3 months ago
- NVV-SuperBench: Beyond Words, Beyond Quality—Benchmarking Nonverbal Vocalizations in Speech Generation (Interspeech 2026 long paper)☆18Jun 21, 2026Updated last month
- OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video☆19Apr 24, 2026Updated 3 months ago
- [IEEE TCSVT] Preprocessing Enhanced Image Compression for Machine Vision☆19Mar 23, 2025Updated last year
- Official repo for Directional Self-supervised Learning for Heavy Image Augmentations [CVPR2022]☆12Jun 29, 2022Updated 4 years ago
- ☆28Jan 9, 2025Updated last year
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Feb 11, 2026Updated 5 months ago
- Build coherent and visually polished multimodal webpages with hierarchical planning, AIGC tools, and iterative reflection.☆15May 17, 2026Updated 2 months ago
- [ECCV 2026 Oral] Official implementation of "OmniForcing: Unleashing Real-time Joint Audio-Visual Generation"[arXiv:2603.11647]. OmniForc…☆173Jul 23, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆11Mar 4, 2025Updated last year
- Code for "OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation"☆151Jun 18, 2026Updated last month
- HippoMM: Hippocampal-inspired Multimodal Memory☆22May 22, 2025Updated last year
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models☆32Feb 3, 2026Updated 6 months ago
- Official Code of NAVA: Native Audio-Visual Alignment for Generation.☆214Jun 30, 2026Updated last month
- ASLP Summer Inter@NPU☆13Jul 30, 2024Updated 2 years ago
- ☆18Jun 13, 2026Updated last month
- [NeurIPS 2024] The official implementation of "Image Copy Detection for Diffusion Models"☆18Oct 1, 2024Updated last year
- [ECCV 2024] RGBD GS-ICP SLAM☆14Nov 5, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- Official respository for ReasonGen-R1☆75Jun 23, 2025Updated last year
- MTVCraft: An Open Veo3-style Audio-Video Generation Demo☆98Oct 8, 2025Updated 9 months ago
- Official code for the ICLR 2025 paper, "Ada-K Routing: Boosting the Efficiency of MoE-based LLMs"☆12Mar 1, 2025Updated last year
- [ICML'26] Toward Human-like Audio-Visual Intelligence of Omni-MLLMs☆17Jun 20, 2026Updated last month
- ☆58Apr 28, 2026Updated 3 months ago
- [ECCV 2026] ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling☆174Jun 23, 2026Updated last month
- Awesome GAN-based Image Restoration☆12Mar 11, 2024Updated 2 years ago
- ☆21Mar 2, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR26] Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs☆33Dec 9, 2025Updated 7 months ago
- LEMAS‑Edit is a multilingual speech editing system, supporting 10 languages: Chinese English Spanish Russian French German Italian Portug…☆21Mar 31, 2026Updated 4 months ago
- ☆12Sep 13, 2024Updated last year
- ☆57Apr 22, 2026Updated 3 months ago
- [NeurIPS 2025 Spotlight] Demo implementation of MoCha Towards Movie-Grade Talking Character Synthesis☆17Dec 27, 2025Updated 7 months ago
- ☆88Nov 16, 2025Updated 8 months ago
- ☆27Sep 10, 2025Updated 10 months ago