Code for "AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs"
β26Oct 9, 2025Updated 10 months ago
Alternatives and similar repositories for AudioMarathon
Users that are interested in AudioMarathon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Adaptive Multimodal Reasoning via Reinforcement Learningβ24Jan 11, 2026Updated 7 months ago
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning (π₯The Exploration of R1 for General Audio-Visβ¦β79Jun 3, 2026Updated 2 months ago
- [CVPR 2026] Variation-aware Vision Token Dropping for Faster Large Vision-Language Modelsβ35May 27, 2026Updated 3 months ago
- Native full-duplex speech dialogue inference for BayLing-Duplex.β80Jun 22, 2026Updated 2 months ago
- The official repository TimeAudio, a comprehensive framework that incorporates fine-grained acoustic cues into LALMs with enhanced moduleβ¦β31Nov 18, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official repo for MMAU-Pro Benchmarkβ22Sep 25, 2025Updated 11 months ago
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.β146Apr 7, 2026Updated 4 months ago
- [APSIPA'22] Exploring Speaker Age Estimation on Different Self-Supervised Learning Modelsβ14Oct 19, 2022Updated 3 years ago
- Multi-Task Speech classification of accent and gender of an english speaker on Mozilla's common voice datasetβ27Jul 17, 2026Updated last month
- β53Jul 5, 2026Updated last month
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness (https://arxiv.org/abs/2511.07931)β84Dec 23, 2025Updated 8 months ago
- β18Feb 1, 2026Updated 6 months ago
- β21Apr 9, 2026Updated 4 months ago
- Open SingSong - Implementation of 'SingSong: Generating Musical Accompaniments from Singing' by Google Research, with a few modificationsβ17Jun 10, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. (CVPR 2025))β62Jun 9, 2025Updated last year
- OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarioβ¦β17Oct 28, 2025Updated 10 months ago
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalitiesβ47Jul 26, 2026Updated last month
- Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Modelβ26May 21, 2026Updated 3 months ago
- GitHub repository for AudioToolAgentβ21Feb 13, 2026Updated 6 months ago
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocolsβ20Nov 19, 2025Updated 9 months ago
- Implementation and experiment of the MusGConv paper.β15Sep 6, 2024Updated last year
- β90Feb 24, 2026Updated 6 months ago
- Native End-to-End Full-Duplex Spoken Language Modelβ128Aug 18, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotationsβ43Oct 15, 2025Updated 10 months ago
- β13Oct 23, 2018Updated 7 years ago
- LibriSpeech-Long is a benchmark dataset for long-form speech generation and processing. Released as part of "Long-Form Speech Generation β¦β99Dec 28, 2024Updated last year
- β18Feb 6, 2026Updated 6 months ago
- Fully Open-source Multimodal Language Models for Science Discoveryβ168Mar 20, 2026Updated 5 months ago
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsβ51Jul 12, 2026Updated last month
- Python scripts to create noisy and reverberant 2-speaker mixture audio with Libri-Light and WHAMβ17Nov 7, 2024Updated last year
- β27Sep 10, 2025Updated 11 months ago
- β28May 22, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β19Mar 2, 2024Updated 2 years ago
- Official Codebase of "A Closer Look at Weakly-Supervised Audio-Visual Source Localization" (NeurIPS 2022)β22Dec 6, 2022Updated 3 years ago
- (ICLR 2026 π₯) Code for "The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs"β79Feb 9, 2026Updated 6 months ago
- Official inference code for UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice.β39May 30, 2026Updated 3 months ago
- [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLMβ27Updated this week
- UTAUTAI(Unrestricted Tune Automated Technology Artificial Interigence)β17Oct 27, 2023Updated 2 years ago
- An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUsβ33Mar 15, 2026Updated 5 months ago