Onset-and-Offset-Aware Sound Event Detection
☆21Feb 10, 2025Updated last year
Alternatives and similar repositories for sed-hsmm
Users that are interested in sed-hsmm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Apr 16, 2026Updated 4 months ago
- ☆146May 13, 2025Updated last year
- ☆29Oct 17, 2024Updated last year
- Official page of "DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis"☆28Apr 15, 2026Updated 4 months ago
- text to speech☆10Mar 19, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Improving Recording Device Generalization using Impulse Response Augmentation☆21Apr 24, 2025Updated last year
- This repository contains the code of the CP JKU submission to DCASE23 Task 1 "Low-complexity Acoustic Scene Classification"☆33Sep 18, 2023Updated 2 years ago
- ☆22Jun 12, 2025Updated last year
- Prediction of sound event bounding boxes (SEBBs)☆35Aug 2, 2024Updated 2 years ago
- Multi-talker ASR based on DiCoW with Serialized Output Training☆21Sep 18, 2025Updated 10 months ago
- ☆33Updated this week
- ☆19Oct 9, 2025Updated 10 months ago
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into one☆26Aug 5, 2024Updated 2 years ago
- (WIP)long form speech generatoins☆30Apr 2, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- This is a repository of neural full-rank spatial covariance analysis with speaker activity (neural FCASA).☆41Mar 12, 2025Updated last year
- ☆23Oct 17, 2024Updated last year
- Unsupervised Voice Activity Detection by Modeling Source and System Information using Zero Frequency Filtering☆23Oct 19, 2023Updated 2 years ago
- Towards a general language-audio model for computational paralinguistic tasks☆31Dec 14, 2024Updated last year
- Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining☆16Oct 12, 2025Updated 10 months ago
- ☆12Nov 7, 2024Updated last year
- The source code for Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation published in IEEE TASLPRO.☆19Jun 3, 2026Updated 2 months ago
- ☆25Jan 24, 2023Updated 3 years ago
- FINALLY: Fast and universal speech enhancement model delivering studio-quality audio for a wide range of recordings.☆28Apr 1, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Export an ONNX graph that performs ISTFT. Designed for TTS models.☆28Apr 23, 2024Updated 2 years ago
- ☆22Sep 14, 2025Updated 11 months ago
- ☆21Mar 6, 2026Updated 5 months ago
- CST-former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection (ICASSP 2024)☆40May 20, 2025Updated last year
- ☆24Mar 19, 2025Updated last year
- ☆10Feb 18, 2022Updated 4 years ago
- Sing any popular song with your voice☆11Jul 10, 2022Updated 4 years ago
- T5Voice is a lightweight PyTorch implementation of T5-based text-to-speech synthesis, supporting both streaming and non-streaming speech …☆28Nov 7, 2025Updated 9 months ago
- This is the official code for ACM CIKM 2025 Paper: ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive …☆60Dec 21, 2025Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [INTERSPEECH 2026] Pre-training, SFT, DPO and GRPO for Text-to-Audio Generation☆49Apr 17, 2026Updated 4 months ago
- Unofficial implementation of ConvNeXt-TTS powered by lightning☆18Oct 20, 2024Updated last year
- [ICASSP 2025] AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder☆14Mar 11, 2025Updated last year
- ☆31Jun 17, 2026Updated 2 months ago
- ☆15May 25, 2026Updated 2 months ago
- PyTorch implementation of WaveFit [2022, Google] which is one of SOTA lightweight/fast speech vocoders.☆71Jul 13, 2026Updated last month
- This repository contains prompts & best practices to annotate audio clips with a very high degree of details using Audio-Language-Models☆35Oct 13, 2024Updated last year