Audio-JEPA is an adaptation of the Joint-Embedding Predictive Architecture (JEPA) for self-supervised audio representation learning. Built upon the I-JEPA paradigm, it uses a Vision Transformer (ViT) backbone to predict latent representations of masked spectrogram patches.
☆65Jul 16, 2026Updated last week
Alternatives and similar repositories for Audio-JEPA
Users that are interested in Audio-JEPA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- JEPAs for audio representation learning☆26Jun 11, 2026Updated last month
- Variable Bitrate Residual Vector Quantization for Audio Coding☆54May 1, 2025Updated last year
- A simple command line tool to calculate WER for ASR.☆14Oct 14, 2024Updated last year
- The MIR-MLPop dataset and the official implementation of the paper "MIR-MLPop: A Multilingual Pop Music Dataset with Time-Aligned Lyrics …☆35Apr 22, 2024Updated 2 years ago
- This is the official codebase for WavJEPA. Time-domain audio foundation model for holistic downstream tasks. "Self-supervised learning fr…☆34Feb 28, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆35Sep 6, 2025Updated 10 months ago
- Forced alignment decoder for Whisper.☆16Mar 13, 2024Updated 2 years ago
- ☆101Jan 19, 2026Updated 6 months ago
- 5Hz Deep-Compression Speech VAE for AR-Diffusion and CALMs☆57Nov 19, 2025Updated 8 months ago
- The official repository for the paper “NonVerbalSpeech-38K: A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understandi…☆68Dec 26, 2025Updated 6 months ago
- Ultra-low-bitrate Speech Codec for Speech Language Modeling Applications☆92Dec 20, 2024Updated last year
- LibriSpeech-Long is a benchmark dataset for long-form speech generation and processing. Released as part of "Long-Form Speech Generation …☆99Dec 28, 2024Updated last year
- This is the accompanying repository to the paper - Automatic Estimation of Singing Voice Musical Dynamics☆16Oct 28, 2024Updated last year
- ☆130Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆16May 7, 2026Updated 2 months ago
- semantic tokenizer for speech and music☆20Jul 6, 2025Updated last year
- CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval☆13Jun 27, 2025Updated last year
- Text-To-Speech for NotebookLM☆39Jul 20, 2025Updated last year
- ☆46Jul 5, 2026Updated 2 weeks ago
- Codebase for the paper 'EncodecMAE: Leveraging neural codecs for universal audio representation learning'☆101Jul 24, 2024Updated 2 years ago
- Joint Embedding Predictive Architecture for Musical Stem Compatibility Estimation☆55Aug 6, 2024Updated last year
- [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix☆214Feb 25, 2026Updated 4 months ago
- Training, validation, and inference code for various SSL approaches and architectures.☆87Apr 7, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pushing the Limits of Zero-shot End-to-End Speech Translation☆25Dec 12, 2024Updated last year
- Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation☆67Jun 16, 2026Updated last month
- EVAR ~ Evaluation package for Audio Representations☆81Feb 19, 2026Updated 5 months ago
- ☆15Apr 16, 2026Updated 3 months ago
- PyTorch implementation of "Source Separation by Flow Matching (FLOSS)" by Google DeepMind☆96Nov 24, 2025Updated 8 months ago
- ☆19Sep 20, 2025Updated 10 months ago
- This is a subset of the DALI set consisting of 240 polyphonic recordings that is used to benchmark lyrics transcription evaluation.☆12Nov 30, 2021Updated 4 years ago
- ☆73Jul 17, 2026Updated last week
- poorman's ar-dit tts☆45Dec 31, 2025Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'☆163Mar 26, 2026Updated 3 months ago
- Inference codebase for "Cacophony: An Improved Contrastive Audio-Text Model". Preprint: https://arxiv.org/abs/2402.06986☆49Jan 19, 2026Updated 6 months ago
- Text-to-text alignment algorithm for speech recognition error analysis.☆32Jun 23, 2026Updated last month
- The inference and trainging code for WordVoice.☆61Jul 17, 2026Updated last week
- A dataset of pitch curves for music performance assessment☆11Jun 5, 2023Updated 3 years ago
- The official repository of SpeechCraft dataset, a large-scale expressive bilingual speech dataset with natural language descriptions.☆197Feb 28, 2026Updated 4 months ago
- Audio-FLAN☆161Sep 23, 2025Updated 10 months ago