The code and weight for LoVA. LoVA is a novel model for Long-form Video-to-Audio generation. Based on the Diffusion Transformer (DiT) architecture, LoVA proves to be more effective at generating long-form audio compared to existing autoregressive models and UNet-based diffusion models.
☆16Feb 27, 2025Updated last year
Alternatives and similar repositories for LoVA
Users that are interested in LoVA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AFlow & MathAI☆18Feb 24, 2025Updated last year
- [ICLR 2026] Evaluating Text Creativity across Diverse Domains: A Dataset and a Large Language Model Evaluator☆18Feb 28, 2026Updated 4 months ago
- [ACM MM 2024] See or Guess: Counterfactually Regularized Image Captioning☆16Feb 17, 2025Updated last year
- 🌟 SwarmAgent: A framework for simulating social group dynamics using multi-agent collaboration, aiding insights into collective behavior…☆13Dec 5, 2023Updated 2 years ago
- ☆22Aug 18, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- An automatic prompt iteration and optimization generator suitable for any scenario☆16Jan 31, 2025Updated last year
- Implementation of an efficient LLM architecture: the Pair-In / Pair-Out Model (PIPO)☆42Jun 10, 2026Updated last month
- TL;DR: We propose a large-scale cross-domain persuasion dataset covers 13,000 scenarios in 35 domains, with the developed PersuGPT model …☆17Feb 12, 2025Updated last year
- [ACM MM 2022] (Oral): Multi-Modal Experience Inspired AI Creation☆21Nov 27, 2024Updated last year
- Source code for "Synchformer: Efficient Synchronization from Sparse Cues" (ICASSP 2024)☆130Sep 15, 2025Updated 10 months ago
- Tools for the evaluation of audio captioning.☆19May 23, 2020Updated 6 years ago
- Implementation of Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching (NeurIPS'24)☆62Apr 3, 2025Updated last year
- 聚合人大:基于知识图谱的高校信息集成与推荐平台开发与应用。2020年中国大学生创新实验计划国家级立项且获得优秀结项。☆12May 28, 2021Updated 5 years ago
- source code for NAACL2022 main conference "Dynamic Programming in Rank Space: Scaling Structured Inference with Low-Rank HMMs and PCFGs"☆10Sep 26, 2022Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- waverless: A serverless framework written by rust with WASM, CRIU, FunctionGraph, Integrated Storage☆13May 22, 2025Updated last year
- PyTorch Implementation of Stepwise Monotonic Multihead Attention similar to Enhancing Monotonicity for Robust Autoregressive Transformer …☆39May 16, 2021Updated 5 years ago
- ☆12Dec 10, 2018Updated 7 years ago
- ☆19Jul 22, 2025Updated last year
- Evaluation of generated videos on the FETV benchmark☆10Apr 6, 2025Updated last year
- ☆28Jul 6, 2026Updated 2 weeks ago
- [EMNLP'2024 Findings] Explore generated documents for enhanced IR with LLMs. We enhance BM25 to surpass strong dense retriever on many da…☆14Mar 28, 2025Updated last year
- Implementation of [CodingGenie: A Proactive LLM-Powered Programming Assistant]☆13Jan 14, 2025Updated last year
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆97Jun 29, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- code for paper Multi-View Joint Graph Representation Learning for Urban Region Embedding. (to be updated)☆34Sep 2, 2021Updated 4 years ago
- Official Code For EMNLP2025 Findings: {DLPO : Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Le…☆10Dec 25, 2025Updated 6 months ago
- ☆17Dec 12, 2023Updated 2 years ago
- 同济大学数据挖掘课程期末作业:股票走势预测☆10Jan 11, 2021Updated 5 years ago
- UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts☆41Jun 12, 2025Updated last year
- A Python runtime for multi-entity AI collaboration — agents, humans, and tools on a shared protocol layer.☆51Jun 18, 2026Updated last month
- 抖音直播网页版弹幕爬取 python 实现☆17Jan 22, 2024Updated 2 years ago
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. (CVPR 2025))☆61Jun 9, 2025Updated last year
- ☆17Nov 20, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆30Jun 30, 2020Updated 6 years ago
- 知予人工智能:从学习者到研究者☆14Jan 20, 2025Updated last year
- [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM☆27Feb 10, 2026Updated 5 months ago
- ☆12Sep 27, 2024Updated last year
- Code for "Memory Efficient Meta-Learning with Large Images"☆11Nov 24, 2021Updated 4 years ago
- [CVPR'26] TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding☆25Jan 4, 2026Updated 6 months ago
- ☆14Apr 25, 2025Updated last year