Code for "AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs"
β26Oct 9, 2025Updated last year
Alternatives and similar repositories for AudioMarathon
Users that are interested in AudioMarathon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Adaptive Multimodal Reasoning via Reinforcement Learningβ24Jan 11, 2026Updated 8 months ago
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning (π₯The Exploration of R1 for General Audio-Visβ¦β80Jun 3, 2026Updated 4 months ago
- [CVPR 2026] Variation-aware Vision Token Dropping for Faster Large Vision-Language Modelsβ36May 27, 2026Updated 4 months ago
- Native full-duplex speech dialogue inference for BayLing-Duplex.β93Jun 22, 2026Updated 3 months ago
- The official repository TimeAudio, a comprehensive framework that incorporates fine-grained acoustic cues into LALMs with enhanced moduleβ¦β31Nov 18, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official repo for MMAU-Pro Benchmarkβ23Sep 25, 2025Updated last year
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.β147Apr 7, 2026Updated 6 months ago
- [APSIPA'22] Exploring Speaker Age Estimation on Different Self-Supervised Learning Modelsβ14Oct 19, 2022Updated 3 years ago
- Multi-Task Speech classification of accent and gender of an english speaker on Mozilla's common voice datasetβ27Jul 17, 2026Updated 2 months ago
- β60Jul 5, 2026Updated 3 months ago
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness (https://arxiv.org/abs/2511.07931)β85Dec 23, 2025Updated 9 months ago
- β18Feb 1, 2026Updated 8 months ago
- Open SingSong - Implementation of 'SingSong: Generating Musical Accompaniments from Singing' by Google Research, with a few modificationsβ17Jun 10, 2024Updated 2 years ago
- β22Apr 9, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. (CVPR 2025))β63Jun 9, 2025Updated last year
- OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarioβ¦β17Oct 28, 2025Updated 11 months ago
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalitiesβ48Jul 26, 2026Updated 2 months ago
- GitHub repository for AudioToolAgentβ21Feb 13, 2026Updated 7 months ago
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocolsβ20Nov 19, 2025Updated 10 months ago
- Implementation and experiment of the MusGConv paper.β15Sep 6, 2024Updated 2 years ago
- β92Feb 24, 2026Updated 7 months ago
- Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Modelβ29May 21, 2026Updated 4 months ago
- Native End-to-End Full-Duplex Spoken Language Modelβ152Aug 18, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotationsβ43Aug 29, 2026Updated last month
- β14Oct 23, 2018Updated 7 years ago
- LibriSpeech-Long is a benchmark dataset for long-form speech generation and processing. Released as part of "Long-Form Speech Generation β¦β99Dec 28, 2024Updated last year
- β20Feb 6, 2026Updated 8 months ago
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsβ51Jul 12, 2026Updated 2 months ago
- Python scripts to create noisy and reverberant 2-speaker mixture audio with Libri-Light and WHAMβ17Nov 7, 2024Updated last year
- β27Sep 10, 2025Updated last year
- β29May 22, 2026Updated 4 months ago
- β19Mar 2, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official Codebase of "A Closer Look at Weakly-Supervised Audio-Visual Source Localization" (NeurIPS 2022)β22Dec 6, 2022Updated 3 years ago
- (ICLR 2026 π₯) Code for "The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs"β79Feb 9, 2026Updated 8 months ago
- [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLMβ27Aug 27, 2026Updated last month
- UTAUTAI(Unrestricted Tune Automated Technology Artificial Interigence)β17Oct 27, 2023Updated 2 years ago
- An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUsβ33Mar 15, 2026Updated 6 months ago
- β29Feb 23, 2026Updated 7 months ago
- Official inference code for UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice.β46May 30, 2026Updated 4 months ago