FastAPI + MLX offline-first voice agent with <1s latency. Minimal UI
☆56Oct 21, 2025Updated 11 months ago
Alternatives and similar repositories for offline-voice-ai
Users that are interested in offline-voice-ai are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 7 months ago
- ☆18May 21, 2026Updated 4 months ago
- Text Match Cut Video Generator Web App☆38Feb 19, 2026Updated 7 months ago
- ClaudeCode to OpenCode Migration guide☆16Jan 17, 2026Updated 8 months ago
- Bloat Free, Portable and Lightweight LLM Frontend (Single HTML file). With Lorebook, Web Search, Macro Engine etc.☆22Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆13Mar 10, 2025Updated last year
- ☆18Mar 17, 2025Updated last year
- From-scratch implementation of OpenAI's GPT-OSS model in Python. No Torch, No GPUs.☆110Nov 5, 2025Updated 10 months ago
- ☆59Feb 8, 2026Updated 7 months ago
- assistant that runs entirely on‑device on Apple‑silicon Macs (M‑series). Chats with a 4‑bit Llama‑3 model accelerated by MLX, and speak…☆16Jun 13, 2025Updated last year
- your private, personal assistant☆81Apr 16, 2026Updated 5 months ago
- Pure C wrapper library to use llama.cpp with Linux and Windows as simple as possible.☆15Sep 6, 2026Updated 3 weeks ago
- 🗣️ Real‑time, low‑latency voice, vision, and conversational‑memory AI assistant built on LiveKit and local LLMs☆110Jun 25, 2025Updated last year
- A markdown web renderer for AI agents — see the web without screenshots☆68Aug 28, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A sophisticated biologically inspired memory system for AI agents. Provides organic, high quality, persistent memory with self-maintenanc…☆73May 31, 2025Updated last year
- Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp☆87Jul 12, 2026Updated 2 months ago
- OpenAI-compatible TTS API that unifies multiple backends with smart chunking for unlimited-length generation☆50Aug 3, 2026Updated last month
- ☆16Feb 24, 2025Updated last year
- Store, query, and create YAML workflow playbooks for LLM agents via MCP. STDIO or Streamable HTTP.☆31Updated this week
- Dashboard v5 Coming Soon!!☆64Feb 15, 2026Updated 7 months ago
- Qwen2-VL for OCR & VQA☆19Sep 3, 2024Updated 2 years ago
- A Multi-Agentic AI Assistant/Builder☆28Jun 1, 2026Updated 3 months ago
- An optimized FastAPI server for OpenAI's Whisper whisper-large-v3-turbo model using MLX optimization☆14Jun 5, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Crack - Make your lid loud!☆16Mar 22, 2026Updated 6 months ago
- A personal Cloudflare-hosted development environment. Work lives in conversations, not terminal tabs.☆18Jun 14, 2026Updated 3 months ago
- Blazing-fast rust implementation of Sesame's Conversational Speech Model (CSM)☆87Mar 26, 2026Updated 6 months ago
- Decentralizing distribution of open-source AI models.☆24Updated this week
- Collection of official scripts created by the Dione Team.☆15Feb 21, 2026Updated 7 months ago
- Simple inbound/outbound packet sniffer☆31Oct 2, 2024Updated last year
- Ragamuffin - Chat with your document, articles or code☆13Nov 13, 2024Updated last year
- ☆14Sep 16, 2024Updated 2 years ago
- Create text chunks which end at natural stopping points without using a tokenizer☆26Nov 26, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆26Aug 26, 2025Updated last year
- Vector functions and indexing for SQLite☆10Mar 26, 2023Updated 3 years ago
- A full GUI experience on top of llama.cpp: all-knobs model tuning, one-click build/update from upstream, HuggingFace discovery with VRAM-…☆88Sep 6, 2026Updated 3 weeks ago
- GPU-accelerated voice assistant — local LLM, fine-tuned Whisper, Kokoro TTS, AMD ROCm☆25Apr 2, 2026Updated 5 months ago
- A production-ready multi-tenant RAG as a Service (RaaS) orchestrator☆31Nov 10, 2025Updated 10 months ago
- MindsApplied EEG Signal Filter for Real-time or Offline Analysis☆17Feb 25, 2026Updated 7 months ago
- Predictive Incident Management analyses large data sets to identify risk patterns, predict outcomes, and guide teams on effective decisio…☆16Nov 25, 2022Updated 3 years ago