☆19Jul 8, 2026Updated 2 months ago
Alternatives and similar repositories for VABench
Users that are interested in VABench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- JoVA: Unified Multimodal Learning for Joint Video-Audio Generation☆34Dec 22, 2025Updated 8 months ago
- [ICML26] AVGen-Bench is a task-driven benchmark for multi-granular evaluation of Text-to-Audio-Video (T2AV) generation.☆31Jul 2, 2026Updated 2 months ago
- The Source Code for MT-Video-Bench @ ACL Findings 2026☆21Jan 20, 2026Updated 8 months ago
- ☆22Nov 25, 2025Updated 9 months ago
- Official code repo for our work "Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models"☆54Jun 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆10Dec 16, 2020Updated 5 years ago
- Launch two rosbot robots in ROS2 for cooperative mapping and navigation☆10Mar 3, 2022Updated 4 years ago
- Code for paper "RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text"☆18May 30, 2024Updated 2 years ago
- Code for "OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation"☆159Jun 18, 2026Updated 3 months ago
- ☆17Mar 24, 2026Updated 5 months ago
- Dataflow-MM, multi-media operators for Dataflow. We aim to prepare data for Multimodal Large Language Models.☆48Apr 13, 2026Updated 5 months ago
- Ego4DSounds: A diverse egocentric dataset with high action-audio correspondence☆21Jun 14, 2024Updated 2 years ago
- 北大软微2023秋季 软工期末必修课真题☆14Jan 21, 2024Updated 2 years ago
- Audio propagation engine - Meta Reality Labs Research.☆24Nov 1, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"☆19Mar 9, 2026Updated 6 months ago
- Implementation of "Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation"☆15Sep 11, 2026Updated last week
- 满分实验:An independent, standalone, and functioning database management system (DBMS) supporting a subset of SQL. Cross-platform. Totally fr…☆14Jun 26, 2021Updated 5 years ago
- ☆19Oct 9, 2025Updated 11 months ago
- Official repository for "Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", https://arxiv.org/abs/2601.1983…☆101Mar 9, 2026Updated 6 months ago
- SMPL/SMPL-H version of HumanML3D☆15Aug 1, 2024Updated 2 years ago
- ☆19Dec 3, 2025Updated 9 months ago
- [NeurIPS 2025] SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models.☆18May 26, 2026Updated 3 months ago
- SQL database written in C++, based on the course Stanford CS 346 and redbase.☆10Jul 12, 2016Updated 10 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code for the paper "DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models"☆17Oct 23, 2025Updated 10 months ago
- The official implementation of MaskGRPO: Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models. (ICLR 2026, arxiv…☆19Jan 27, 2026Updated 7 months ago
- ☆20Aug 23, 2024Updated 2 years ago
- ☆16May 30, 2025Updated last year
- [ECCV 2026 Oral] Official implementation of "OmniForcing: Unleashing Real-time Joint Audio-Visual Generation"[arXiv:2603.11647]. OmniForc…☆195Jul 23, 2026Updated last month
- Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill☆25Jun 8, 2026Updated 3 months ago
- Vim plugin to copy text to Windows clipboard on WSL☆12Jan 8, 2023Updated 3 years ago
- Table logger using Rich☆13Aug 13, 2025Updated last year
- ☆13Jul 14, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆12Sep 23, 2024Updated last year
- Code for paper "Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy" [NeurIPS 2025] .☆18Dec 6, 2025Updated 9 months ago
- ☆12Mar 22, 2025Updated last year
- A Benchmark Corpus for Low-Resource Cantonese Punctuation Restoration from Speech Transcripts☆15Dec 3, 2024Updated last year
- Deep Research as Rubric for Reinforcement Learning☆25Jun 30, 2026Updated 2 months ago
- [ACL 2023] To Copy Rather Than Memorize: A Vertical Learning Paradigm for Knowledge Graph Completion☆11Feb 3, 2023Updated 3 years ago
- This repository documents Barry's journey in learning deep learning for speech processing. Here, you'll find scripts and code snippets re…☆13Oct 8, 2025Updated 11 months ago