[ICML26] AVGen-Bench is a task-driven benchmark for multi-granular evaluation of Text-to-Audio-Video (T2AV) generation.
☆31Jul 2, 2026Updated 3 months ago
Alternatives and similar repositories for AVGen-Bench
Users that are interested in AVGen-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆24Apr 24, 2026Updated 5 months ago
- CastleHill: Separable Causal Diffusion / Varitaion Flow Maps for LTX-2 long-form video generation☆15May 19, 2026Updated 4 months ago
- A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation☆34Jun 9, 2026Updated 3 months ago
- ☆20Jul 8, 2026Updated 2 months ago
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆48Jul 26, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 🧂 [ECCV 2026] Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation☆21Aug 24, 2026Updated last month
- NVV-SuperBench: Beyond Words, Beyond Quality—Benchmarking Nonverbal Vocalizations in Speech Generation (Interspeech 2026 long paper)☆19Jun 21, 2026Updated 3 months ago
- OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video☆22Apr 24, 2026Updated 5 months ago
- [IEEE TCSVT] Preprocessing Enhanced Image Compression for Machine Vision☆19Mar 23, 2025Updated last year
- Official repo for Directional Self-supervised Learning for Heavy Image Augmentations [CVPR2022]☆12Jun 29, 2022Updated 4 years ago
- ☆28Jan 9, 2025Updated last year
- D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning☆15Aug 4, 2026Updated last month
- Build coherent and visually polished multimodal webpages with hierarchical planning, AIGC tools, and iterative reflection.☆18May 17, 2026Updated 4 months ago
- [ECCV 2026 Oral] Official implementation of "OmniForcing: Unleashing Real-time Joint Audio-Visual Generation"[arXiv:2603.11647]. OmniForc…☆197Jul 23, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆11Mar 4, 2025Updated last year
- CodexOpt: Optize your Agents.MD and Skills for Codex with GEPA☆21May 26, 2026Updated 4 months ago
- Code for "OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation"☆157Jun 18, 2026Updated 3 months ago
- Multi-Faceted Distillation of Base-Novel Commonality for Few-shot Object Detection, ECCV 2022☆35Mar 23, 2023Updated 3 years ago
- [ICML 2025] Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions☆16Mar 7, 2026Updated 6 months ago
- HippoMM: Hippocampal-inspired Multimodal Memory☆21May 22, 2025Updated last year
- Astra planner and JEV controller for Minecraft, with native recording, tested routes, and run verification.☆571Sep 20, 2026Updated last week
- [NeurIPS 2026] Official Code of NAVA: Native Audio-Visual Alignment for Generation.☆226Jun 30, 2026Updated 3 months ago
- ASLP Summer Inter@NPU☆13Jul 30, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is a ComfyUI plugin for https://github.com/OpenMOSS/MOVA☆22Jan 30, 2026Updated 8 months ago
- ☆21Jun 13, 2026Updated 3 months ago
- ☆18Oct 5, 2024Updated last year
- [NeurIPS 2024] The official implementation of "Image Copy Detection for Diffusion Models"☆18Oct 1, 2024Updated 2 years ago
- [ECCV 2024] RGBD GS-ICP SLAM☆14Nov 5, 2024Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- 天池竞赛☆10Sep 12, 2016Updated 10 years ago
- Official respository for ReasonGen-R1☆74Jun 23, 2025Updated last year
- [CVPR 2023] Adversarial Robustness via Random Projection Filters☆13Jun 20, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR2026] Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception☆19Jan 26, 2026Updated 8 months ago
- [ICML'26] Toward Human-like Audio-Visual Intelligence of Omni-MLLMs☆18Updated this week
- [ECCV 2026] ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling☆185Sep 16, 2026Updated 2 weeks ago
- ☆58Apr 28, 2026Updated 5 months ago
- Awesome GAN-based Image Restoration☆12Mar 11, 2024Updated 2 years ago
- LEMAS‑Edit is a multilingual speech editing system, supporting 10 languages: Chinese English Spanish Russian French German Italian Portug…☆22Mar 31, 2026Updated 6 months ago
- [NeurIPS 2025 Spotlight] Demo implementation of MoCha Towards Movie-Grade Talking Character Synthesis☆20Dec 27, 2025Updated 9 months ago