A high-performance batch audio transcription tool using nvidia/parakeet-tdt-0.6b-v2 to generate accurate, well-segmented SRT subtitles, with special optimizations for very long audio files.
☆18Dec 9, 2025Updated 8 months ago
Alternatives and similar repositories for parakeet-tdt-0.6b-v2-Batch-Transcriber
Users that are interested in parakeet-tdt-0.6b-v2-Batch-Transcriber are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pure-PyTorch Parakeet TDT inference☆52Mar 10, 2026Updated 4 months ago
- SLT 2024 Challenge: Post-ASR-Speaker-Tagging☆16Jun 16, 2024Updated 2 years ago
- ☆24Aug 1, 2026Updated last week
- SLISEMAP: Combining supervised dimensionality reduction with local explanations☆21Apr 24, 2025Updated last year
- ☆34Jul 21, 2026Updated 2 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 🤯 A code collection for learning and exploration.☆18Mar 14, 2025Updated last year
- A demo-level low-latency, high-throughput inference engine for whisper☆20Nov 9, 2025Updated 9 months ago
- ☆19Aug 22, 2025Updated 11 months ago
- PyTorch implementation of RWKV blocks☆31Jul 22, 2025Updated last year
- Extract video from Samsung Motion Photo. Supports JPEG, HEIF/HEIC☆20Dec 11, 2025Updated 7 months ago
- ☆13Apr 16, 2025Updated last year
- AAAI-26 Oral | Implementation of "LoKI: Low-damage Knowledge Implanting of Large Language Models"☆22Apr 11, 2026Updated 3 months ago
- Ultra-Sortformer for Scalable Speaker Diarization☆27Apr 9, 2026Updated 4 months ago
- Compute WER and SER for speech recognition evaluation☆28Jun 6, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- We implemented the DEMUCS model for speech enhancement in the time-frequency domain, and additionally implemented HD-DEMUCS.☆34Nov 8, 2023Updated 2 years ago
- A repo containing download guidance and corresponding scripts of the VoxBlink dataset.☆31Apr 16, 2024Updated 2 years ago
- Finetune Nemo parakeet ASR model with new language (support 8 bit optimizer). Experimental birwkv-fastconformer TDT for long-form ASR(8.5…☆26Nov 27, 2025Updated 8 months ago
- A comprehensive manager and fixer for Gemini CLI (@google/gemini-cli) and Qwen CLI (@qwen-code/qwen-code) on Termux (Android).☆16Apr 25, 2026Updated 3 months ago
- LLM-based ASR recipe with Zipformer encoder and Qwen LLM☆35Sep 25, 2025Updated 10 months ago
- A unified evaluation suite for speech-to-text translation, covering SpeechLLMs, SFMs, and cascaded systems across diverse real-world spee…☆32Updated this week
- ☆37Jan 6, 2026Updated 7 months ago
- Turn SMS Backup & Restore SMS XML files into HTML transcripts☆27May 17, 2019Updated 7 years ago
- An onnx-exportable implementation of iSTFT in torch☆35Feb 19, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- RWKV Batch infer backend ⚡Base on albatross https://github.com/BlinkDL/Albatross 🕊️☆37Jul 27, 2026Updated 2 weeks ago
- Project video on my Youtube channel about building an audio content analyzer dashboard.☆23Feb 22, 2023Updated 3 years ago
- Eureka-Audio: A 1.7B lightweight audio–language model that matches 7B–30B models on ASR, audio understanding, and paralinguistic reasonin…☆42Apr 11, 2026Updated 3 months ago
- Official Repository for "Efficient Vocal Source Separation Through Windowed RoFormer"☆46Oct 30, 2025Updated 9 months ago
- Convert your ChatGPT Message History into Data Viz and Markdown Notes☆29Oct 26, 2023Updated 2 years ago
- wav2vec2 audio classification for prosodic boundary detection and other tasks☆42Aug 11, 2023Updated 2 years ago
- Variations of L1 SNR Loss function for training audio source separation machine learning models☆45Updated this week
- Softened ROSA QKV Operators for Training Next-Generation LLM Models☆39Updated this week
- This repository implement a novel zero-shot TTS framework, named Flamed-TTS, focusing on the efficient generation and dynamic pacing in …☆57Aug 9, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official repository for Mamba-based Segmentation Model for Speaker Diarization☆47May 13, 2025Updated last year
- A web-based application to parse, view, and manage SMS backup files (XML format) from "SMS Backup & Restore" with advanced features like …☆29Aug 2, 2026Updated last week
- Implementation of 2-simplicial attention proposed by Clift et al. (2019) and the recent attempt to make practical in Fast and Simplex, Ro…☆49Sep 2, 2025Updated 11 months ago
- Export the STFT or ISTFT process in ONNX format.☆47Jun 6, 2026Updated 2 months ago
- Gemma-based Multilingual Machine Translation Models☆53Updated this week
- Faster whisper Running on AMD GPUs with modified CTranslate 2 Libraries served up with Wyoming protocol☆34Aug 17, 2024Updated last year
- sms-db is a tool to build an SQLite database out of collections of SMS and MMS messages in various formats. The database can then be quer…☆45Dec 7, 2021Updated 4 years ago