Audio captioning - DCASE challenge 2023 task 6a
☆30Dec 26, 2024Updated last year
Alternatives and similar repositories for audio-captioning
Users that are interested in audio-captioning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CoNeTTE: An efficient Audio Captioning system leveraging multiple datasets with Task Embedding☆23Dec 17, 2025Updated 7 months ago
- [NCMMSC]☆16Feb 19, 2025Updated last year
- Evaluate EfficientAT models on the Holistic Evaluation of Audio Representations Benchmark.☆34Jun 23, 2023Updated 3 years ago
- Unsupervised spoken sentence embeddings☆14Dec 14, 2022Updated 3 years ago
- ☆16Feb 10, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Scripts for computing common lyrics-to-audio alignment evaluation metrics. Usable evaluation for any token-based alignment (e.g. if tok…☆18Oct 27, 2020Updated 5 years ago
- ☆10Aug 29, 2024Updated last year
- Official repository of https://doi.org/10.1109/TASLP.2022.3167258. More up-to-date code is in "refactor" branch.☆192Jun 8, 2023Updated 3 years ago
- Simple LPC vocoder in Python☆13Jan 7, 2022Updated 4 years ago
- ☆15Sep 26, 2022Updated 3 years ago
- Phoneme Level Lyrics Alignment and Text-Informed Singing Voice Separation☆24Nov 8, 2021Updated 4 years ago
- Zerospeech Challenge 2021: validation and evaluation software☆12Jun 13, 2022Updated 4 years ago
- Audio captioning recipe☆53Oct 23, 2025Updated 9 months ago
- (Hybrid) BYOL-S feature extractor using serab-byols package in pytorch.☆27Apr 20, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5☆19Nov 29, 2022Updated 3 years ago
- Open SingSong - Implementation of 'SingSong: Generating Musical Accompaniments from Singing' by Google Research, with a few modifications☆17Jun 10, 2024Updated 2 years ago
- Code repository for the paper "Improving End-to-End SLU performance with Prosodic Attention and Distillation" accepted at Interspeech 202…☆27May 17, 2023Updated 3 years ago
- ☆13Jun 30, 2026Updated last month
- Use your hands to control a Music Player☆12Oct 25, 2024Updated last year
- LLM-Codec: Neural Audio Codec Meets Language Model Objectives☆23May 3, 2026Updated 3 months ago
- DSing ASR task: Resources and Baseline for an unaccompanied singing ASR.☆19Jul 9, 2026Updated last month
- ☆33Jul 18, 2024Updated 2 years ago
- CUDA-accelerated PyTorch implementation of t-SNE☆25May 15, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- In Divisive we have all points in one cluster initially and we break the cluster into required number of clusters.☆10May 19, 2018Updated 8 years ago
- Python code for handling the Clotho dataset.☆85Nov 24, 2020Updated 5 years ago
- Learning graph embedding using DeepWalk or Node2Vec☆15Jan 3, 2017Updated 9 years ago
- Python and C/C++ library for fast, accurate PCA on the GPU☆12Jun 4, 2018Updated 8 years ago
- Tensorflow implementation of VQVAE for voice conversion☆12Apr 3, 2018Updated 8 years ago
- ☆10Sep 25, 2024Updated last year
- NIST SPH File reader (e.g. for TEDLIUM Corpus)☆26May 2, 2020Updated 6 years ago
- Code for the method proposed in the paper:- ccc-wav2vec 2.0: Clustering aided Cross-Contrastive learning of Self-Supervised speech repres…☆23Mar 18, 2024Updated 2 years ago
- ☆31Feb 4, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆22Sep 26, 2022Updated 3 years ago
- Word Discovery in Visually Grounded, Self-Supervised Speech Models☆27Dec 4, 2023Updated 2 years ago
- ACE Studio PC版工程文件解密工具☆13Jul 13, 2022Updated 4 years ago
- DRFI For Region Dissection☆13Jan 11, 2019Updated 7 years ago
- Language independent SSL-based Speaker Anonymization system☆20May 28, 2024Updated 2 years ago
- ☆10Feb 17, 2017Updated 9 years ago
- simple trainer for musicgen/audiocraft☆15Jul 14, 2023Updated 3 years ago