A variable-frame-rate 16 kHz speech codec based on FocalCodec
☆21Feb 11, 2026Updated 7 months ago
Alternatives and similar repositories for dycast
Users that are interested in dycast are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A low-bitrate single-codebook 16 / 24 kHz speech codec based on focal modulation☆176Nov 30, 2025Updated 10 months ago
- A collections of audio codecs with a standardized API☆45Apr 15, 2026Updated 5 months ago
- [ICLR2026] FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates☆54Aug 20, 2026Updated last month
- Codebase for 'ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining'☆27Updated this week
- Llama-Mimi is a speech language model that uses a unified tokenizer (Mimi) and a single Transformer decoder (Llama) to jointly model sequ…☆31Sep 20, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Official code for paper:"Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding"☆39Jan 28, 2026Updated 8 months ago
- Voxtral Codec : Combining Semantic VQ and Acoustic FSQ for Ultra-Low Bitrate Speech Generation (Voxtral TTS Backbone)☆17Mar 27, 2026Updated 6 months ago
- A method that directly addresses the modality gap by aligning speech token with the corresponding text transcription during the tokenizat…☆123Sep 3, 2025Updated last year
- [ICASSP 2026]Official code for "Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum"☆27Jan 22, 2026Updated 8 months ago
- ☆31Feb 14, 2026Updated 7 months ago
- Extract a target speaker’s clean, non-overlapped speech from multi-speaker audio and export word-safe LJSpeech-style TTS datasets.☆23Jun 14, 2026Updated 3 months ago
- ICASSP 2024 - Generative De-Quantization for Neural Speech Codec via Latent Diffusion.☆56Jul 31, 2026Updated 2 months ago
- ☆14Aug 1, 2025Updated last year
- Bayesian deep learning for remaining useful life estimation via Stein variational gradient descent☆30Feb 5, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Kanade is a single-layer disentangled speech tokenizer that extracts compact tokens suitable for both generative and discriminative model…☆118Jul 18, 2026Updated 2 months ago
- Neural Speech Codec☆26Jan 25, 2021Updated 5 years ago
- Official implementation of "Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech G…☆16Feb 6, 2026Updated 7 months ago
- The source code for Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation published in IEEE TASLPRO.☆19Jun 3, 2026Updated 4 months ago
- For IEEE ASRU(2025)☆15Jun 21, 2025Updated last year
- Detect individual instruments activity in an audio file. 🎤🎹🎸🥁☆18Jun 29, 2021Updated 5 years ago
- ☆36Mar 29, 2025Updated last year
- ☆26Mar 1, 2026Updated 7 months ago
- Open-weights voice acting pipeline combining zero-shot voice cloning with natural-language direction. Provide a reference voice (or gener…☆18May 25, 2026Updated 4 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- FastWave is a lightweight diffusion model for general audio super-resolution (any -> 48 kHz). SOTA quality reconstruction metrics with ju…☆21May 16, 2026Updated 4 months ago
- Variations of L1 SNR Loss function for training audio source separation machine learning models☆45Sep 1, 2026Updated last month
- A family of efficient speech models for multilingual phone recognition☆84Updated this week
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into one☆26Aug 5, 2024Updated 2 years ago
- libvits-ncnn is an ncnn implementation of the VITS library that enables cross-platform GPU-accelerated speech synthesis.🎙️💻☆62May 6, 2023Updated 3 years ago
- Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding☆19May 5, 2025Updated last year
- ☆14Oct 3, 2025Updated last year
- Official implementation of the paper "BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec"☆221Sep 19, 2024Updated 2 years ago
- A neural speech codec based on discrete WavLM representations☆27Aug 28, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- TensorFlow,DCGAN,VAE,LSTM,CNN,Acoustic Scene Classification☆11Jun 5, 2019Updated 7 years ago
- fd-sds☆21Apr 8, 2026Updated 5 months ago
- PyTorch code implementation of EfficientSpeech - to be presented at ICASSP2023.☆183Mar 18, 2024Updated 2 years ago
- A highly optimized engine for neutts-air model to generate minutes of audio in seconds. Over 200x realtime on modern hardware!☆119Nov 24, 2025Updated 10 months ago
- ☆19Jul 23, 2025Updated last year
- High fidelity neural audio codec for TTS models☆37Dec 22, 2025Updated 9 months ago
- Extract phoneme-level timestamps from speeh audio.☆167Sep 22, 2026Updated last week