A high-performance batch audio transcription tool using nvidia/parakeet-tdt-0.6b-v2 to generate accurate, well-segmented SRT subtitles, with special optimizations for very long audio files.
☆18Dec 9, 2025Updated 7 months ago
Alternatives and similar repositories for parakeet-tdt-0.6b-v2-Batch-Transcriber
Users that are interested in parakeet-tdt-0.6b-v2-Batch-Transcriber are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pure-PyTorch Parakeet TDT inference☆48Mar 10, 2026Updated 4 months ago
- SLT 2024 Challenge: Post-ASR-Speaker-Tagging☆16Jun 16, 2024Updated 2 years ago
- 🤯 A code collection for learning and exploration.☆18Mar 14, 2025Updated last year
- ☆23Jul 12, 2026Updated last week
- SLISEMAP: Combining supervised dimensionality reduction with local explanations☆21Apr 24, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A demo-level low-latency, high-throughput inference engine for whisper☆20Nov 9, 2025Updated 8 months ago
- ☆19Aug 22, 2025Updated 10 months ago
- PyTorch implementation of RWKV blocks☆31Jul 22, 2025Updated 11 months ago
- Extract video from Samsung Motion Photo. Supports JPEG, HEIF/HEIC☆20Dec 11, 2025Updated 7 months ago
- ☆13Apr 16, 2025Updated last year
- AAAI-26 Oral | Implementation of "LoKI: Low-damage Knowledge Implanting of Large Language Models"☆22Apr 11, 2026Updated 3 months ago
- Ultra-Sortformer for Scalable Speaker Diarization☆27Apr 9, 2026Updated 3 months ago
- Compute WER and SER for speech recognition evaluation☆27Jun 6, 2026Updated last month
- We implemented the DEMUCS model for speech enhancement in the time-frequency domain, and additionally implemented HD-DEMUCS.☆34Nov 8, 2023Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A repo containing download guidance and corresponding scripts of the VoxBlink dataset.☆30Apr 16, 2024Updated 2 years ago
- Finetune Nemo parakeet ASR model with new language (support 8 bit optimizer). Experimental birwkv-fastconformer TDT for long-form ASR(8.5…☆26Nov 27, 2025Updated 7 months ago
- A comprehensive manager and fixer for Gemini CLI (@google/gemini-cli) and Qwen CLI (@qwen-code/qwen-code) on Termux (Android).☆16Apr 25, 2026Updated 2 months ago
- LLM-based ASR recipe with Zipformer encoder and Qwen LLM☆34Sep 25, 2025Updated 9 months ago
- ☆37Jan 6, 2026Updated 6 months ago
- A unified evaluation suite for speech-to-text translation, covering SpeechLLMs, SFMs, and cascaded systems across diverse real-world spee…☆32Apr 25, 2026Updated 2 months ago
- Turn SMS Backup & Restore SMS XML files into HTML transcripts☆27May 17, 2019Updated 7 years ago
- An onnx-exportable implementation of iSTFT in torch☆35Feb 19, 2025Updated last year
- RWKV Batch infer backend ⚡Base on albatross https://github.com/BlinkDL/Albatross 🕊️☆36Jul 13, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Project video on my Youtube channel about building an audio content analyzer dashboard.☆23Feb 22, 2023Updated 3 years ago
- Eureka-Audio: A 1.7B lightweight audio–language model that matches 7B–30B models on ASR, audio understanding, and paralinguistic reasonin…☆40Apr 11, 2026Updated 3 months ago
- Official Repository for "Efficient Vocal Source Separation Through Windowed RoFormer"☆45Oct 30, 2025Updated 8 months ago
- Convert your ChatGPT Message History into Data Viz and Markdown Notes☆27Oct 26, 2023Updated 2 years ago
- Native End-to-End Full-Duplex Spoken Language Model☆54Updated this week
- wav2vec2 audio classification for prosodic boundary detection and other tasks☆42Aug 11, 2023Updated 2 years ago
- Variations of L1 SNR Loss function for training audio source separation machine learning models☆45May 1, 2026Updated 2 months ago
- Softened ROSA QKV Operators for Training Next-Generation LLM Models☆39Jun 26, 2026Updated 3 weeks ago
- This repository implement a novel zero-shot TTS framework, named Flamed-TTS, focusing on the efficient generation and dynamic pacing in …☆57Aug 9, 2025Updated 11 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- Implementation of 2-simplicial attention proposed by Clift et al. (2019) and the recent attempt to make practical in Fast and Simplex, Ro…☆49Sep 2, 2025Updated 10 months ago
- A web-based application to parse, view, and manage SMS backup files (XML format) from "SMS Backup & Restore" with advanced features like …☆26Jun 4, 2025Updated last year
- Export the STFT or ISTFT process in ONNX format.☆47Jun 6, 2026Updated last month
- Gemma-based Multilingual Machine Translation Models☆51Feb 13, 2026Updated 5 months ago
- Faster whisper Running on AMD GPUs with modified CTranslate 2 Libraries served up with Wyoming protocol☆34Aug 17, 2024Updated last year
- sms-db is a tool to build an SQLite database out of collections of SMS and MMS messages in various formats. The database can then be quer…☆45Dec 7, 2021Updated 4 years ago