☆21Jun 4, 2026Updated last month
Alternatives and similar repositories for S2S-Arena
Users that are interested in S2S-Arena are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols☆20Nov 19, 2025Updated 8 months ago
- Chinese Prosodic Structure Prediction☆10May 18, 2019Updated 7 years ago
- Open repository of "MSU-Bench: Towards Understanding the Conversational Multi-Speaker Scenarios"☆17Jul 7, 2026Updated 2 weeks ago
- ☆18Sep 22, 2024Updated last year
- a fully open-source implementation of a GPT-4o-like speech-to-speech video understanding model.☆38Apr 7, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆14Oct 3, 2025Updated 9 months ago
- ☆17Jun 10, 2026Updated last month
- The official Soundwave repository☆223Mar 16, 2025Updated last year
- Please visit https://thuhcsi.github.io/SnakeGAN/☆37Apr 25, 2023Updated 3 years ago
- ☆33May 19, 2026Updated 2 months ago
- ☆18Mar 6, 2026Updated 4 months ago
- We propose C2SER, a novel audio-language model designed to enhance the stability and accuracy of speech emotion recognition through conte…☆49Mar 3, 2025Updated last year
- An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs☆33Mar 15, 2026Updated 4 months ago
- Demo for AudioSAE paper☆15Apr 26, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment☆14Feb 5, 2025Updated last year
- [Findings of NAACL 2024] Source code of paper CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers a…☆68Mar 31, 2024Updated 2 years ago
- ☆17Apr 9, 2026Updated 3 months ago
- The "GPT-API-Accelerate" project provides a set of Python classes for accelerating the process of generating responses to prompts using t…☆23Oct 12, 2024Updated last year
- This repository includes the full release of PrinciplismQA dataset and assessment scripts.☆16Updated this week
- Official Implementation and Dataset of paper - DFADD: The Diffusion and Flow-matching based Audio Deepfake Dataset☆16Apr 7, 2025Updated last year
- GitHub repository for AudioToolAgent☆20Feb 13, 2026Updated 5 months ago
- ☆11Dec 24, 2024Updated last year
- ☆21Sep 14, 2025Updated 10 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- This repository contains source codes for SoftCTC. Original paper can be found here: https://arxiv.org/abs/2212.02135☆19Mar 7, 2023Updated 3 years ago
- Onset-and-Offset-Aware Sound Event Detection☆21Feb 10, 2025Updated last year
- egrecho project☆11Apr 30, 2026Updated 2 months ago
- Multi-talker ASR based on DiCoW with Serialized Output Training☆20Sep 18, 2025Updated 10 months ago
- Official code for "WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling"☆62Jun 27, 2026Updated 3 weeks ago
- ☆15Oct 24, 2025Updated 8 months ago
- ☆61Oct 28, 2024Updated last year
- Montreal Forced Aligner for Vietnamese☆15Oct 23, 2023Updated 2 years ago
- Official repository for U-SAM (Interspeech 2025)☆28Jun 3, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Adaptive Multimodal Reasoning via Reinforcement Learning☆23Jan 11, 2026Updated 6 months ago
- Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models☆55Sep 2, 2025Updated 10 months ago
- [Interspeech 2024] LiteFocus is a tool designed to accelerate diffusion-based TTA model, now implemented with the base model AudioLDM2.☆34Mar 11, 2025Updated last year
- ☆15Jul 24, 2025Updated 11 months ago
- ☆33Nov 27, 2021Updated 4 years ago
- end-to-end text to audio scene generation model☆50Jun 16, 2026Updated last month
- ☆123May 18, 2026Updated 2 months ago