Audio captioning - DCASE challenge 2023 task 6a
☆30Dec 26, 2024Updated last year
Alternatives and similar repositories for audio-captioning
Users that are interested in audio-captioning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CoNeTTE: An efficient Audio Captioning system leveraging multiple datasets with Task Embedding☆23Dec 17, 2025Updated 9 months ago
- Evaluate EfficientAT models on the Holistic Evaluation of Audio Representations Benchmark.☆34Jun 23, 2023Updated 3 years ago
- Unsupervised spoken sentence embeddings☆14Dec 14, 2022Updated 3 years ago
- ☆16Feb 10, 2026Updated 7 months ago
- Scripts for computing common lyrics-to-audio alignment evaluation metrics. Usable evaluation for any token-based alignment (e.g. if tok…☆18Oct 27, 2020Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆10Aug 29, 2024Updated 2 years ago
- Official repository of https://doi.org/10.1109/TASLP.2022.3167258. More up-to-date code is in "refactor" branch.☆192Jun 8, 2023Updated 3 years ago
- Simple LPC vocoder in Python☆13Jan 7, 2022Updated 4 years ago
- ☆15Sep 26, 2022Updated 3 years ago
- Phoneme Level Lyrics Alignment and Text-Informed Singing Voice Separation☆24Nov 8, 2021Updated 4 years ago
- Zerospeech Challenge 2021: validation and evaluation software☆12Jun 13, 2022Updated 4 years ago
- (Hybrid) BYOL-S feature extractor using serab-byols package in pytorch.☆27Apr 20, 2024Updated 2 years ago
- ASR text preprocessing utility☆21Aug 5, 2024Updated 2 years ago
- ☆10Oct 25, 2019Updated 6 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆12May 23, 2023Updated 3 years ago
- Code for T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5☆19Nov 29, 2022Updated 3 years ago
- Open SingSong - Implementation of 'SingSong: Generating Musical Accompaniments from Singing' by Google Research, with a few modifications☆17Jun 10, 2024Updated 2 years ago
- Optimizing speaker verification and spoofing countermeasure systems together with REINFORCE☆13Mar 31, 2021Updated 5 years ago
- Code repository for the paper "Improving End-to-End SLU performance with Prosodic Attention and Distillation" accepted at Interspeech 202…☆27May 17, 2023Updated 3 years ago
- ☆12Nov 17, 2015Updated 10 years ago
- ☆13Jun 30, 2026Updated 2 months ago
- Use your hands to control a Music Player☆12Oct 25, 2024Updated last year
- LLM-Codec: Neural Audio Codec Meets Language Model Objectives☆23May 3, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- DSing ASR task: Resources and Baseline for an unaccompanied singing ASR.☆19Jul 9, 2026Updated 2 months ago
- ☆33Jul 18, 2024Updated 2 years ago
- In Divisive we have all points in one cluster initially and we break the cluster into required number of clusters.☆10May 19, 2018Updated 8 years ago
- Python code for handling the Clotho dataset.☆85Nov 24, 2020Updated 5 years ago
- Tensorflow implementation of VQVAE for voice conversion☆12Apr 3, 2018Updated 8 years ago
- Code for the method proposed in the paper:- ccc-wav2vec 2.0: Clustering aided Cross-Contrastive learning of Self-Supervised speech repres…☆23Mar 18, 2024Updated 2 years ago
- ☆31Feb 4, 2021Updated 5 years ago
- ☆22Sep 26, 2022Updated 3 years ago
- Word Discovery in Visually Grounded, Self-Supervised Speech Models☆27Dec 4, 2023Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ACE Studio PC版工程文件解密工具☆13Jul 13, 2022Updated 4 years ago
- Language independent SSL-based Speaker Anonymization system☆20May 28, 2024Updated 2 years ago
- ☆35Apr 10, 2023Updated 3 years ago
- This repository contains the source code for the implementation of two deep learning models concerning the audio super resolution task.☆14Mar 14, 2023Updated 3 years ago
- Prosodic Speech Segmentation with Transformers☆28Feb 25, 2024Updated 2 years ago
- Official repository for the paper "AudioMAE++: learning better masked audio representations with SwiGLU FFNs"☆16Apr 30, 2026Updated 4 months ago
- Tools for the evaluation of audio captioning.☆19May 23, 2020Updated 6 years ago