☆21Jun 4, 2026Updated 2 months ago
Alternatives and similar repositories for S2S-Arena
Users that are interested in S2S-Arena are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols☆20Nov 19, 2025Updated 8 months ago
- [ICLR 2026] StableToken: A state-of-the-art noise-robust semantic speech tokenizer featuring Voting-LFQ for resilient SpeechLLMs.☆33Feb 27, 2026Updated 5 months ago
- ☆18Sep 22, 2024Updated last year
- ☆33Nov 4, 2025Updated 9 months ago
- Please visit https://thuhcsi.github.io/SnakeGAN/☆37Apr 25, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Chinese Prosodic Structure Prediction☆10May 18, 2019Updated 7 years ago
- ☆21Sep 14, 2025Updated 10 months ago
- We propose C2SER, a novel audio-language model designed to enhance the stability and accuracy of speech emotion recognition through conte…☆50Mar 3, 2025Updated last year
- The official Soundwave repository☆223Mar 16, 2025Updated last year
- Repository containing the website for the EMNLP 2023 conference☆17Feb 12, 2025Updated last year
- ☆24Sep 20, 2024Updated last year
- Open repository of "MSU-Bench: Towards Understanding the Conversational Multi-Speaker Scenarios"☆20Jul 7, 2026Updated last month
- a fully open-source implementation of a GPT-4o-like speech-to-speech video understanding model.☆38Apr 7, 2025Updated last year
- ☆14Oct 3, 2025Updated 10 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆61Oct 28, 2024Updated last year
- Llama-Mimi is a speech language model that uses a unified tokenizer (Mimi) and a single Transformer decoder (Llama) to jointly model sequ…☆31Sep 20, 2025Updated 10 months ago
- ☆25Jul 22, 2026Updated 2 weeks ago
- AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data☆36Dec 31, 2023Updated 2 years ago
- ☆33Updated this week
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.☆145Apr 7, 2026Updated 4 months ago
- Evaluation tool used in the BigVSAN paper☆14Mar 22, 2024Updated 2 years ago
- ☆36May 19, 2026Updated 2 months ago
- Towards Fine-grained Audio Captioning with Multimodal Contextual Cues☆88Jan 4, 2026Updated 7 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models☆26Aug 11, 2024Updated last year
- ☆20Mar 6, 2026Updated 5 months ago
- ☆28Aug 1, 2026Updated last week
- Implementation of the paper "Variable Bitrate Residual Vector Quantization for Audio Coding"☆11Apr 10, 2025Updated last year
- Official implementation of TISDiSS, a scalable framework for discriminative source separation.☆16Jul 31, 2026Updated last week
- An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs☆33Mar 15, 2026Updated 4 months ago
- Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment☆14Feb 5, 2025Updated last year
- Demo for AudioSAE paper☆15Apr 26, 2026Updated 3 months ago
- [Findings of NAACL 2024] Source code of paper CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers a…☆68Mar 31, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆19Apr 9, 2026Updated 4 months ago
- This repository includes the full release of PrinciplismQA dataset and assessment scripts.☆16Jul 14, 2026Updated 3 weeks ago
- Official Implementation and Dataset of paper - DFADD: The Diffusion and Flow-matching based Audio Deepfake Dataset☆16Apr 7, 2025Updated last year
- The official repository for the paper “NonVerbalSpeech-38K: A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understandi…☆68Dec 26, 2025Updated 7 months ago
- GitHub repository for AudioToolAgent☆20Feb 13, 2026Updated 5 months ago
- ☆60Apr 1, 2026Updated 4 months ago
- A rigorous framework for evaluating and guiding the development of next-generation AI assistants.☆19Jan 26, 2026Updated 6 months ago