☆28Jul 17, 2026Updated 2 weeks ago
Alternatives and similar repositories for f-actor
Users that are interested in f-actor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Implementation of NAACL 2025 Paper: Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models☆18Apr 30, 2025Updated last year
- fd-sds☆21Apr 8, 2026Updated 3 months ago
- ☆32Updated this week
- Native full-duplex speech dialogue inference for BayLing-Duplex.☆64Jun 22, 2026Updated last month
- Codebase for 'ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining'☆24Jun 20, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL 2025] OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching☆45Feb 9, 2025Updated last year
- Multi-talker ASR based on DiCoW with Serialized Output Training☆21Sep 18, 2025Updated 10 months ago
- [ACL 2026 Main] Training, inference, and testing of the SAC speech codec model.☆109Nov 1, 2025Updated 9 months ago
- We propose C2SER, a novel audio-language model designed to enhance the stability and accuracy of speech emotion recognition through conte…☆50Mar 3, 2025Updated last year
- Llama-Mimi is a speech language model that uses a unified tokenizer (Mimi) and a single Transformer decoder (Llama) to jointly model sequ…☆31Sep 20, 2025Updated 10 months ago
- Understanding and Tackling Hallucinations in Large Audio-Language Models | ICASSP 2025, Interspeech 2024☆34Mar 14, 2025Updated last year
- [ICLR2026] FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates☆51Jul 1, 2026Updated last month
- [ICLR 2026] StableToken: A state-of-the-art noise-robust semantic speech tokenizer featuring Voting-LFQ for resilient SpeechLLMs.☆33Feb 27, 2026Updated 5 months ago
- FLM-Audio is a audio-language subversion of RoboEgo/FLM-Ego -- an omnimodal model with native full duplexity.☆70May 15, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The official repository TimeAudio, a comprehensive framework that incorporates fine-grained acoustic cues into LALMs with enhanced module…☆30Nov 18, 2025Updated 8 months ago
- [ICASSP 2026] Official code for "Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration"☆17Apr 16, 2026Updated 3 months ago
- Towards a general language-audio model for computational paralinguistic tasks☆31Dec 14, 2024Updated last year
- Implementation of "Look, Listen and Recognise:character-aware audio-visual subtitling"☆21Nov 3, 2025Updated 9 months ago
- ☆24Sep 20, 2024Updated last year
- ☆89Feb 24, 2026Updated 5 months ago
- ☆20Mar 6, 2026Updated 5 months ago
- Official Pytorch implementation of "Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models" [IEEE ICASSP 202…☆38Mar 10, 2026Updated 4 months ago
- MichiAI: A Low Latency, Full Duplex Speech LLM with zero coherence loss☆111Apr 24, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- LibriSpeech-Long is a benchmark dataset for long-form speech generation and processing. Released as part of "Long-Form Speech Generation …☆99Dec 28, 2024Updated last year
- A Benchmark for Evaluating Turn-Taking and Overlap Handling in Full-Duplex Spoken Dialogue Models☆253May 20, 2026Updated 2 months ago
- ☆74Jul 29, 2026Updated last week
- DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action☆118May 20, 2026Updated 2 months ago
- Unofficial fairseq-free PyTorch implementation of UTMOS (v1, 2022), matching the original system.☆35Jun 6, 2026Updated 2 months ago
- Text-To-Speech for NotebookLM☆39Jul 20, 2025Updated last year
- Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis☆27Mar 21, 2025Updated last year
- Demo for AudioSAE paper☆15Apr 26, 2026Updated 3 months ago
- [ICLR 2026] Official implementation of Toward Complex-Valued Neural Networks for Waveform Generation☆20Apr 10, 2026Updated 3 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆19Apr 9, 2026Updated 3 months ago
- ☆18Feb 1, 2026Updated 6 months ago
- Eureka-Audio: A 1.7B lightweight audio–language model that matches 7B–30B models on ASR, audio understanding, and paralinguistic reasonin…☆41Apr 11, 2026Updated 3 months ago
- Extract phoneme-level timestamps from speeh audio.☆158Jun 7, 2026Updated last month
- ☆48Jul 25, 2026Updated last week
- [INTERSPEECH 2025] Official code for "SEED: Speaker Embedding Enhancement Diffusion Model"☆59Nov 3, 2025Updated 9 months ago
- Curated list for papers, codes and resources related to Text-to-Audio (TTA) Generation☆76Jul 20, 2026Updated 2 weeks ago