Code for "AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs"
β26Oct 9, 2025Updated 10 months ago
Alternatives and similar repositories for AudioMarathon
Users that are interested in AudioMarathon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Adaptive Multimodal Reasoning via Reinforcement Learningβ23Jan 11, 2026Updated 6 months ago
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning (π₯The Exploration of R1 for General Audio-Visβ¦β78Jun 3, 2026Updated 2 months ago
- [CVPR 2026] Variation-aware Vision Token Dropping for Faster Large Vision-Language Modelsβ34May 27, 2026Updated 2 months ago
- Native full-duplex speech dialogue inference for BayLing-Duplex.β64Jun 22, 2026Updated last month
- The official repository TimeAudio, a comprehensive framework that incorporates fine-grained acoustic cues into LALMs with enhanced moduleβ¦β30Nov 18, 2025Updated 8 months ago
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Official repo for MMAU-Pro Benchmarkβ22Sep 25, 2025Updated 10 months ago
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.β145Apr 7, 2026Updated 4 months ago
- [APSIPA'22] Exploring Speaker Age Estimation on Different Self-Supervised Learning Modelsβ14Oct 19, 2022Updated 3 years ago
- OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarioβ¦β17Oct 28, 2025Updated 9 months ago
- Multi-Task Speech classification of accent and gender of an english speaker on Mozilla's common voice datasetβ28Jul 17, 2026Updated 3 weeks ago
- β49Jul 5, 2026Updated last month
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness (https://arxiv.org/abs/2511.07931)β81Dec 23, 2025Updated 7 months ago
- β18Feb 1, 2026Updated 6 months ago
- β19Apr 9, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Open SingSong - Implementation of 'SingSong: Generating Musical Accompaniments from Singing' by Google Research, with a few modificationsβ17Jun 10, 2024Updated 2 years ago
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. (CVPR 2025))β61Jun 9, 2025Updated last year
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalitiesβ47Jul 26, 2026Updated 2 weeks ago
- Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Modelβ25May 21, 2026Updated 2 months ago
- GitHub repository for AudioToolAgentβ20Feb 13, 2026Updated 5 months ago
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocolsβ20Nov 19, 2025Updated 8 months ago
- Implementation and experiment of the MusGConv paper.β15Sep 6, 2024Updated last year
- β89Feb 24, 2026Updated 5 months ago
- Native End-to-End Full-Duplex Spoken Language Modelβ100Jul 30, 2026Updated last week
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotationsβ43Oct 15, 2025Updated 9 months ago
- β13Oct 23, 2018Updated 7 years ago
- LibriSpeech-Long is a benchmark dataset for long-form speech generation and processing. Released as part of "Long-Form Speech Generation β¦β99Dec 28, 2024Updated last year
- β18Feb 6, 2026Updated 6 months ago
- Fully Open-source Multimodal Language Models for Science Discoveryβ167Mar 20, 2026Updated 4 months ago
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsβ51Jul 12, 2026Updated 3 weeks ago
- Python scripts to create noisy and reverberant 2-speaker mixture audio with Libri-Light and WHAMβ17Nov 7, 2024Updated last year
- β27Sep 10, 2025Updated 11 months ago
- β28May 22, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- β19Mar 2, 2024Updated 2 years ago
- Official Codebase of "A Closer Look at Weakly-Supervised Audio-Visual Source Localization" (NeurIPS 2022)β22Dec 6, 2022Updated 3 years ago
- (ICLR 2026 π₯) Code for "The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs"β80Feb 9, 2026Updated 6 months ago
- Official inference code for UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice.β33May 30, 2026Updated 2 months ago
- [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLMβ27Feb 10, 2026Updated 6 months ago
- UTAUTAI(Unrestricted Tune Automated Technology Artificial Interigence)β17Oct 27, 2023Updated 2 years ago
- An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUsβ33Mar 15, 2026Updated 4 months ago