Onset-and-Offset-Aware Sound Event Detection
☆21Feb 10, 2025Updated last year
Alternatives and similar repositories for sed-hsmm
Users that are interested in sed-hsmm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Apr 16, 2026Updated 3 months ago
- ☆145May 13, 2025Updated last year
- ☆29Oct 17, 2024Updated last year
- Official page of "DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis"☆26Apr 15, 2026Updated 3 months ago
- text to speech☆10Mar 19, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Improving Recording Device Generalization using Impulse Response Augmentation☆21Apr 24, 2025Updated last year
- This repository contains the code of the CP JKU submission to DCASE23 Task 1 "Low-complexity Acoustic Scene Classification"☆32Sep 18, 2023Updated 2 years ago
- ☆22Jun 12, 2025Updated last year
- Prediction of sound event bounding boxes (SEBBs)☆35Aug 2, 2024Updated last year
- Multi-talker ASR based on DiCoW with Serialized Output Training☆21Sep 18, 2025Updated 10 months ago
- ☆32May 18, 2026Updated 2 months ago
- ☆19Oct 9, 2025Updated 9 months ago
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into one☆26Aug 5, 2024Updated last year
- (WIP)long form speech generatoins☆30Apr 2, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- This is a repository of neural full-rank spatial covariance analysis with speaker activity (neural FCASA).☆40Mar 12, 2025Updated last year
- ☆23Oct 17, 2024Updated last year
- Unsupervised Voice Activity Detection by Modeling Source and System Information using Zero Frequency Filtering☆23Oct 19, 2023Updated 2 years ago
- Towards a general language-audio model for computational paralinguistic tasks☆30Dec 14, 2024Updated last year
- Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining☆16Oct 12, 2025Updated 9 months ago
- ☆12Nov 7, 2024Updated last year
- The source code for Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation published in IEEE TASLPRO.☆18Jun 3, 2026Updated last month
- ☆25Jan 24, 2023Updated 3 years ago
- FINALLY: Fast and universal speech enhancement model delivering studio-quality audio for a wide range of recordings.☆28Apr 1, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Export an ONNX graph that performs ISTFT. Designed for TTS models.☆28Apr 23, 2024Updated 2 years ago
- ☆21Sep 14, 2025Updated 10 months ago
- ☆18Mar 6, 2026Updated 4 months ago
- CST-former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection (ICASSP 2024)☆39May 20, 2025Updated last year
- ☆24Mar 19, 2025Updated last year
- ☆10Feb 18, 2022Updated 4 years ago
- Sing any popular song with your voice☆11Jul 10, 2022Updated 4 years ago
- T5Voice is a lightweight PyTorch implementation of T5-based text-to-speech synthesis, supporting both streaming and non-streaming speech …☆28Nov 7, 2025Updated 8 months ago
- This is the official code for ACM CIKM 2025 Paper: ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive …☆59Dec 21, 2025Updated 7 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [INTERSPEECH 2026] Pre-training, SFT, DPO and GRPO for Text-to-Audio Generation☆48Apr 17, 2026Updated 3 months ago
- Unofficial implementation of ConvNeXt-TTS powered by lightning☆18Oct 20, 2024Updated last year
- [ICASSP 2025] AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder☆14Mar 11, 2025Updated last year
- ☆28Jun 17, 2026Updated last month
- ☆15May 25, 2026Updated 2 months ago
- PyTorch implementation of WaveFit [2022, Google] which is one of SOTA lightweight/fast speech vocoders.☆70Jul 13, 2026Updated 2 weeks ago
- This repository contains prompts & best practices to annotate audio clips with a very high degree of details using Audio-Language-Models☆35Oct 13, 2024Updated last year