☆18Jul 8, 2026Updated last month
Alternatives and similar repositories for VABench
Users that are interested in VABench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML26] AVGen-Bench is a task-driven benchmark for multi-granular evaluation of Text-to-Audio-Video (T2AV) generation.☆30Jul 2, 2026Updated last month
- The Source Code for MT-Video-Bench @ ACL Findings 2026☆21Jan 20, 2026Updated 7 months ago
- ☆20Nov 25, 2025Updated 9 months ago
- JoVA: Unified Multimodal Learning for Joint Video-Audio Generation☆33Dec 22, 2025Updated 8 months ago
- Official code repo for our work "Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models"☆54Jun 17, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Wan: Open and Advanced Large-Scale Video Generative Models☆31Jul 28, 2025Updated last year
- KVAE tokenizers☆73Aug 14, 2026Updated 2 weeks ago
- [NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”☆38Jul 24, 2025Updated last year
- The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.☆39Updated this week
- Code for "OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation"☆156Jun 18, 2026Updated 2 months ago
- ☆17Mar 24, 2026Updated 5 months ago
- Dataflow-MM, multi-media operators for Dataflow. We aim to prepare data for Multimodal Large Language Models.☆49Apr 13, 2026Updated 4 months ago
- Code and dataset release for "PACS: A Dataset for Physical Audiovisual CommonSense Reasoning" (ECCV 2022)☆18Dec 20, 2022Updated 3 years ago
- Ego4DSounds: A diverse egocentric dataset with high action-audio correspondence☆21Jun 14, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 北大软微2023秋季 软工期末必修课真题☆14Jan 21, 2024Updated 2 years ago
- Audio propagation engine - Meta Reality Labs Research.☆24Nov 1, 2022Updated 3 years ago
- [CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"☆19Mar 9, 2026Updated 5 months ago
- Implementation of "Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation"☆15Oct 31, 2024Updated last year
- MTVCraft: An Open Veo3-style Audio-Video Generation Demo☆99Oct 8, 2025Updated 10 months ago
- This tools read and converts a dbf file to a txt file, diferent from a CSV this is simple a data table dump. It is possible to filter an…☆14Apr 6, 2025Updated last year
- Official repository for "Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", https://arxiv.org/abs/2601.1983…☆100Mar 9, 2026Updated 5 months ago
- ☆19Dec 3, 2025Updated 8 months ago
- [NeurIPS 2025] SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models.☆18May 26, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Code for the paper "DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models"☆17Oct 23, 2025Updated 10 months ago
- The official implementation of MaskGRPO: Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models. (ICLR 2026, arxiv…☆19Jan 27, 2026Updated 7 months ago
- ☆20Aug 23, 2024Updated 2 years ago
- A longitudinal dataset for academic literature, including papers, metadata, and citation graphs, Also available on 🤗 HuggingFace and Kag…☆18Sep 6, 2025Updated 11 months ago
- ☆16May 30, 2025Updated last year
- [ECCV 2026 Oral] Official implementation of "OmniForcing: Unleashing Real-time Joint Audio-Visual Generation"[arXiv:2603.11647]. OmniForc…☆191Jul 23, 2026Updated last month
- Vim plugin to copy text to Windows clipboard on WSL☆12Jan 8, 2023Updated 3 years ago
- Table logger using Rich☆13Aug 13, 2025Updated last year
- ☆13Jul 14, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆12Sep 23, 2024Updated last year
- AAAI 2024-Controllable Mind Visual Diffusion Model☆16Dec 18, 2023Updated 2 years ago
- arxiv翻译修复器!☆22Nov 13, 2024Updated last year
- Code for paper "Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy" [NeurIPS 2025] .☆18Dec 6, 2025Updated 8 months ago
- JAX implementation of the Mistral 7b v0.1 model☆13Mar 27, 2024Updated 2 years ago
- ☆12Mar 22, 2025Updated last year
- A Benchmark Corpus for Low-Resource Cantonese Punctuation Restoration from Speech Transcripts☆15Dec 3, 2024Updated last year