[ICML26] AVGen-Bench is a task-driven benchmark for multi-granular evaluation of Text-to-Audio-Video (T2AV) generation.
☆31Jul 2, 2026Updated 2 months ago
Alternatives and similar repositories for AVGen-Bench
Users that are interested in AVGen-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆22Apr 24, 2026Updated 4 months ago
- CastleHill: Separable Causal Diffusion / Varitaion Flow Maps for LTX-2 long-form video generation☆15May 19, 2026Updated 3 months ago
- A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation☆34Jun 9, 2026Updated 3 months ago
- ☆19Jul 8, 2026Updated 2 months ago
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆47Jul 26, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 🧂 [ECCV 2026] Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation☆21Aug 24, 2026Updated 2 weeks ago
- NVV-SuperBench: Beyond Words, Beyond Quality—Benchmarking Nonverbal Vocalizations in Speech Generation (Interspeech 2026 long paper)☆19Jun 21, 2026Updated 2 months ago
- OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video☆21Apr 24, 2026Updated 4 months ago
- [IEEE TCSVT] Preprocessing Enhanced Image Compression for Machine Vision☆19Mar 23, 2025Updated last year
- Official repo for Directional Self-supervised Learning for Heavy Image Augmentations [CVPR2022]☆12Jun 29, 2022Updated 4 years ago
- ☆28Jan 9, 2025Updated last year
- [ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.☆382Aug 30, 2026Updated 2 weeks ago
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Aug 4, 2026Updated last month
- [ECCV 2026 Oral] Official implementation of "OmniForcing: Unleashing Real-time Joint Audio-Visual Generation"[arXiv:2603.11647]. OmniForc…☆194Jul 23, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆11Mar 4, 2025Updated last year
- CodexOpt: Optize your Agents.MD and Skills for Codex with GEPA☆19May 26, 2026Updated 3 months ago
- Code for "OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation"☆158Jun 18, 2026Updated 2 months ago
- Multi-Faceted Distillation of Base-Novel Commonality for Few-shot Object Detection, ECCV 2022☆35Mar 23, 2023Updated 3 years ago
- [ICML 2025] Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions☆16Mar 7, 2026Updated 6 months ago
- HippoMM: Hippocampal-inspired Multimodal Memory☆22May 22, 2025Updated last year
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models☆32Feb 3, 2026Updated 7 months ago
- Official Code of NAVA: Native Audio-Visual Alignment for Generation.☆226Jun 30, 2026Updated 2 months ago
- ASLP Summer Inter@NPU☆13Jul 30, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This is a ComfyUI plugin for https://github.com/OpenMOSS/MOVA☆22Jan 30, 2026Updated 7 months ago
- ☆21Jun 13, 2026Updated 3 months ago
- [NeurIPS 2024] The official implementation of "Image Copy Detection for Diffusion Models"☆18Oct 1, 2024Updated last year
- [ECCV 2024] RGBD GS-ICP SLAM☆14Nov 5, 2024Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- This respository is used for time reasoning task for mult-session dialogue system.☆18Feb 7, 2026Updated 7 months ago
- This project is the official implementation of "UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating"☆93Jun 25, 2026Updated 2 months ago
- Official respository for ReasonGen-R1☆74Jun 23, 2025Updated last year
- [CVPR 2023] Adversarial Robustness via Random Projection Filters☆13Jun 20, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICLR2026] Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception☆18Jan 26, 2026Updated 7 months ago
- MTVCraft: An Open Veo3-style Audio-Video Generation Demo☆99Oct 8, 2025Updated 11 months ago
- 🏠 [ECCV 2024] The core gsplat component for GaussianImage☆42Mar 13, 2026Updated 6 months ago
- [ICML'26] Toward Human-like Audio-Visual Intelligence of Omni-MLLMs☆18Jun 20, 2026Updated 2 months ago
- [CVPR 2025] T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation☆123Oct 25, 2025Updated 10 months ago
- ☆23Mar 2, 2026Updated 6 months ago
- An imaginary extension of rotary position embeddings for long-context language models☆33Dec 9, 2025Updated 9 months ago