Code for "AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs"
β26Oct 9, 2025Updated 11 months ago
Alternatives and similar repositories for AudioMarathon
Users that are interested in AudioMarathon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Adaptive Multimodal Reasoning via Reinforcement Learningβ24Jan 11, 2026Updated 8 months ago
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning (π₯The Exploration of R1 for General Audio-Visβ¦β80Jun 3, 2026Updated 3 months ago
- [CVPR 2026] Variation-aware Vision Token Dropping for Faster Large Vision-Language Modelsβ35May 27, 2026Updated 3 months ago
- Native full-duplex speech dialogue inference for BayLing-Duplex.β90Jun 22, 2026Updated 2 months ago
- The official repository TimeAudio, a comprehensive framework that incorporates fine-grained acoustic cues into LALMs with enhanced moduleβ¦β31Nov 18, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official repo for MMAU-Pro Benchmarkβ22Sep 25, 2025Updated 11 months ago
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.β147Apr 7, 2026Updated 5 months ago
- [APSIPA'22] Exploring Speaker Age Estimation on Different Self-Supervised Learning Modelsβ14Oct 19, 2022Updated 3 years ago
- Multi-Task Speech classification of accent and gender of an english speaker on Mozilla's common voice datasetβ27Jul 17, 2026Updated 2 months ago
- β60Jul 5, 2026Updated 2 months ago
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness (https://arxiv.org/abs/2511.07931)β84Dec 23, 2025Updated 8 months ago
- β18Feb 1, 2026Updated 7 months ago
- Open SingSong - Implementation of 'SingSong: Generating Musical Accompaniments from Singing' by Google Research, with a few modificationsβ17Jun 10, 2024Updated 2 years ago
- β21Apr 9, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. (CVPR 2025))β62Jun 9, 2025Updated last year
- OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarioβ¦β17Oct 28, 2025Updated 10 months ago
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalitiesβ48Jul 26, 2026Updated last month
- GitHub repository for AudioToolAgentβ21Feb 13, 2026Updated 7 months ago
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocolsβ20Nov 19, 2025Updated 10 months ago
- Implementation and experiment of the MusGConv paper.β15Sep 6, 2024Updated 2 years ago
- β91Feb 24, 2026Updated 6 months ago
- Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Modelβ27May 21, 2026Updated 3 months ago
- Native End-to-End Full-Duplex Spoken Language Modelβ145Aug 18, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotationsβ43Aug 29, 2026Updated 3 weeks ago
- β14Oct 23, 2018Updated 7 years ago
- LibriSpeech-Long is a benchmark dataset for long-form speech generation and processing. Released as part of "Long-Form Speech Generation β¦β99Dec 28, 2024Updated last year
- β20Feb 6, 2026Updated 7 months ago
- Fully Open-source Multimodal Language Models for Science Discoveryβ171Mar 20, 2026Updated 5 months ago
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsβ51Jul 12, 2026Updated 2 months ago
- Python scripts to create noisy and reverberant 2-speaker mixture audio with Libri-Light and WHAMβ17Nov 7, 2024Updated last year
- β27Sep 10, 2025Updated last year
- β29May 22, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β19Mar 2, 2024Updated 2 years ago
- Official Codebase of "A Closer Look at Weakly-Supervised Audio-Visual Source Localization" (NeurIPS 2022)β22Dec 6, 2022Updated 3 years ago
- (ICLR 2026 π₯) Code for "The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs"β79Feb 9, 2026Updated 7 months ago
- [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLMβ27Aug 27, 2026Updated 3 weeks ago
- UTAUTAI(Unrestricted Tune Automated Technology Artificial Interigence)β17Oct 27, 2023Updated 2 years ago
- An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUsβ33Mar 15, 2026Updated 6 months ago
- β29Feb 23, 2026Updated 6 months ago