Onset-and-Offset-Aware Sound Event Detection
☆22Feb 10, 2025Updated last year
Alternatives and similar repositories for sed-hsmm
Users that are interested in sed-hsmm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Apr 16, 2026Updated 4 months ago
- ☆147May 13, 2025Updated last year
- ☆29Oct 17, 2024Updated last year
- Official page of "DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis"☆31Apr 15, 2026Updated 4 months ago
- text to speech☆10Mar 19, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Improving Recording Device Generalization using Impulse Response Augmentation☆21Apr 24, 2025Updated last year
- This repository contains the code of the CP JKU submission to DCASE23 Task 1 "Low-complexity Acoustic Scene Classification"☆33Sep 18, 2023Updated 2 years ago
- ☆22Jun 12, 2025Updated last year
- Prediction of sound event bounding boxes (SEBBs)☆35Aug 2, 2024Updated 2 years ago
- Multi-talker ASR based on DiCoW with Serialized Output Training☆21Sep 18, 2025Updated 11 months ago
- ☆35Aug 24, 2026Updated last week
- ☆19Oct 9, 2025Updated 10 months ago
- My hybrid TTS network that combines, VALL-E, VoiceBox, SpeechFlow, Seamless and TortoiseTTS into one☆26Aug 5, 2024Updated 2 years ago
- (WIP)long form speech generatoins☆30Apr 2, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is a repository of neural full-rank spatial covariance analysis with speaker activity (neural FCASA).☆41Mar 12, 2025Updated last year
- ☆23Oct 17, 2024Updated last year
- Unsupervised Voice Activity Detection by Modeling Source and System Information using Zero Frequency Filtering☆23Oct 19, 2023Updated 2 years ago
- Towards a general language-audio model for computational paralinguistic tasks☆31Dec 14, 2024Updated last year
- Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining☆16Oct 12, 2025Updated 10 months ago
- ☆12Nov 7, 2024Updated last year
- The source code for Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation published in IEEE TASLPRO.☆19Jun 3, 2026Updated 3 months ago
- ☆25Jan 24, 2023Updated 3 years ago
- FINALLY: Fast and universal speech enhancement model delivering studio-quality audio for a wide range of recordings.☆29Apr 1, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Export an ONNX graph that performs ISTFT. Designed for TTS models.☆28Apr 23, 2024Updated 2 years ago
- ☆23Sep 14, 2025Updated 11 months ago
- ☆21Mar 6, 2026Updated 6 months ago
- CST-former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection (ICASSP 2024)☆40May 20, 2025Updated last year
- ☆23Mar 19, 2025Updated last year
- ☆10Feb 18, 2022Updated 4 years ago
- Sing any popular song with your voice☆11Jul 10, 2022Updated 4 years ago
- T5Voice is a lightweight PyTorch implementation of T5-based text-to-speech synthesis, supporting both streaming and non-streaming speech …☆28Nov 7, 2025Updated 9 months ago
- This is the official code for ACM CIKM 2025 Paper: ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive …☆60Dec 21, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [INTERSPEECH 2026] Pre-training, SFT, DPO and GRPO for Text-to-Audio Generation☆50Apr 17, 2026Updated 4 months ago
- Unofficial implementation of ConvNeXt-TTS powered by lightning☆18Oct 20, 2024Updated last year
- [ICASSP 2025] AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder☆15Mar 11, 2025Updated last year
- ☆34Jun 17, 2026Updated 2 months ago
- ☆15May 25, 2026Updated 3 months ago
- PyTorch implementation of WaveFit [2022, Google] which is one of SOTA lightweight/fast speech vocoders.☆71Jul 13, 2026Updated last month
- This repository aims to collect Transformer-based sound event detection (SED) algorithms.☆104Feb 10, 2026Updated 6 months ago