Ichigo Whisper is a compact (22M parameters), open-source speech tokenizer for the Whisper-medium, designed to enhance performance on multilingual with minimal impact on its original English capabilities.
☆17Jan 20, 2025Updated last year
Alternatives and similar repositories for WhisperSpeech
Users that are interested in WhisperSpeech are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆22Mar 25, 2025Updated last year
- ☆15Updated this week
- 3D Sound Source Localization using Masked Autoencoders☆22Feb 12, 2025Updated last year
- Code release for "TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices"☆23Jun 7, 2025Updated last year
- cortex.llamacpp is a high-efficiency C++ inference engine for edge computing. It is a dynamic library that can be loaded by any server a…☆44Jul 4, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation (INTERSPEECH 2022)☆26Jun 5, 2025Updated last year
- SDK for accessing Vernier sensors via the Go! USB interfaces☆12Sep 25, 2025Updated last year
- App to search images with Unsplash's API and react-query 🔋☆10Oct 7, 2022Updated 3 years ago
- Telegram bot to help you with your findings 🚀☆23Jul 11, 2024Updated 2 years ago
- Vistral-V: Visual Instruction Tuning for Vistral - Vietnamese Large Vision-Language Model.☆23Jul 1, 2024Updated 2 years ago
- This is the sample application for the DevOps Capstone Project on Learn to Cloud. It shows your favorite TV Shows or movies, the front-en…☆20Oct 1, 2024Updated last year
- ☆20Jan 2, 2026Updated 8 months ago
- HiFi-SR is a Python-based pipeline for the detection of plant mitochondrial structural rearrangements based on the mapping of PacBio high…☆13Jun 26, 2026Updated 3 months ago
- The OpenAI Whisper speech-to-text model as a simple HTTP server☆14Oct 26, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Efficient Finetuning for OpenAI GPT-OSS☆24Oct 2, 2025Updated 11 months ago
- Implementation for the different ML tasks on Kaggle platform with GPUs.☆24Jul 23, 2026Updated 2 months ago
- Official repo for the Vietnam-Celeb dataset☆26Aug 27, 2023Updated 3 years ago
- Large Language Models (LLMs) Learning Resources☆21Jun 16, 2024Updated 2 years ago
- A Deep-learning project utilizing 3D human pose estimation to compare different poses☆14Feb 25, 2024Updated 2 years ago
- Anomalous sound detection with machine learning and deep learning☆14Jun 24, 2024Updated 2 years ago
- ☆12Jun 2, 2022Updated 4 years ago
- Deep based Visual SLAM Project(Depth estimation, Optical flow, Visual inertial odometry)☆16Jan 7, 2026Updated 8 months ago
- ☆15Sep 5, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- LIGHTVOC AN UPSAMPLING-FREE GAN VOCODER BASED ON CONFORMER AND INVERSE SHORT-TIME FOURIER TRANSFORM☆18May 17, 2024Updated 2 years ago
- [APSIPA'22] Exploring Speaker Age Estimation on Different Self-Supervised Learning Models☆14Oct 19, 2022Updated 3 years ago
- Datasets used in the paper "Reward hacking behavior can generalize across tasks"☆16Aug 17, 2025Updated last year
- openvino version of openai/whisper☆16Jun 19, 2026Updated 3 months ago
- Visualization tools for audio-only and multi-modal speaker diarization dataset☆13Oct 27, 2023Updated 2 years ago
- Repository for the paper "ViHateT5: Enhancing Hate Speech Detection in Vietnamese with A Unified Text-to-Text Transformer Model" (ACL'202…☆15Aug 13, 2024Updated 2 years ago
- A YOLO-based light weight object tracking system using ByteTrack☆16Jan 17, 2025Updated last year
- noise reduction☆17Jul 3, 2024Updated 2 years ago
- Whisper-Flamingo [Interspeech 2024] and mWhisper-Flamingo [IEEE SPL 2025] for Audio-Visual Speech Recognition and Translation☆209Jul 29, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- StyleTTS2 + Vocos as a Decoder☆13Jul 31, 2026Updated last month
- Low-latency ASR using SpeechBrain StreamingASR and torchaudio StreamReader.☆19Apr 19, 2025Updated last year
- Qwen3-ASR speech-to-text for llama.cpp — patch, GGUF models, and benchmarks☆18Feb 2, 2026Updated 7 months ago
- Text Analysis: Implementation of ULMFiT by Howard & Ruder on Twitter dataset☆10Feb 7, 2019Updated 7 years ago
- ☆124Aug 4, 2026Updated last month
- 🐍 pymecab-ko. you can find original version here: https://bitbucket.org/eunjeon/mecab-ko, https://github.com/SamuraiT/mecab-python3☆25Sep 23, 2025Updated last year
- This project aims to predict smartphone prices using a combination of batch and stream processing techniques in a Big Data environment. T…☆27Apr 15, 2024Updated 2 years ago