TurboQuant KV Cache Compression for llama.cpp — 5.2x memory reduction with near-lossless quality | Implementation of Google DeepMind's TurboQuant (ICLR 2026)
☆92Aug 8, 2026Updated 3 weeks ago
Alternatives and similar repositories for TurboQuant
Users that are interested in TurboQuant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Open-source platform for creating, distributing and running sovereign EU-compliant LLMs. Verticalize any model for your domain, language …☆57Updated this week
- Extract a target speaker’s clean, non-overlapped speech from multi-speaker audio and export word-safe LJSpeech-style TTS datasets.☆23Jun 14, 2026Updated 2 months ago
- Compressed KV cache as cross-backend wire format for Metal + CUDA split inference over Thunderbolt 5☆16Apr 14, 2026Updated 4 months ago
- LLM inference in C/C++☆35Updated this week
- TurboQuant llama.cpp fork with optimized turbo4 kernels for Gemma 4 D=256/512 heads — lazy K/V, batch decode, warp-cooperative write. 120…☆36Apr 5, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Build AI agents using the same architecture patterns as Claude Code. Skill + 6 examples + runnable Python code. No framework.☆23Apr 3, 2026Updated 4 months ago
- LLM inference in C/C++☆2,332Updated this week
- ☆15Jun 23, 2024Updated 2 years ago
- “A locally hosted, memory-aware AI microservice—designed for cultural continuity, decentralized intelligence, and ethical autonomy.”☆27May 1, 2025Updated last year
- NEXO runtime core for NEXO Desktop: local memory, automation, MCP tools and update-managed runtime.☆27Jul 3, 2026Updated last month
- Implementation of different noise embeddings for noise aware training of Kaldi acoustic models.☆13Feb 13, 2021Updated 5 years ago
- 🧠 Bio-Agent OS: 🇻🇳 Bio-Inspired Memory Framework for AI Agents (OpenClaw/ERP). Researched & Developed by Dev Tuan Anh Ha (Locaith Solu…☆20Aug 20, 2026Updated last week
- Embed, cluster, and visualize any collection of texts in 3D semantic space — then learn a continuous semantic flow field over that space,…☆18May 9, 2026Updated 3 months ago
- ☆12Jun 5, 2018Updated 8 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆39Oct 1, 2023Updated 2 years ago
- Distributed workflow automation engine built for AI-native workloads.☆18Updated this week
- ACE-Step: A Step Towards Music Generation Foundation Model☆50May 20, 2025Updated last year
- Updated folk of g2pk☆14Aug 18, 2023Updated 3 years ago
- ☆14Aug 1, 2025Updated last year
- High-performance late-interaction retrieval engine for on-prem AI. ColBERT/ColPali multi-vector search with Rust fused MaxSim, Triton GPU…☆17Jul 6, 2026Updated last month
- Agent Skills for InsForge☆36Updated this week
- Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (…☆26Aug 20, 2026Updated last week
- Simple inference for Vits2 TTS Using ONNXRUNTIME and espeak-ng on C++☆19Apr 17, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆7,013Jul 20, 2026Updated last month
- ⚡ One API for 20+ LLM providers. OpenAI-compatible, single binary, runs forever.☆21Jun 1, 2026Updated 2 months ago
- Fully automated memory and context management for Claude Code using hooks - Zero friction, zero context loss☆33Oct 22, 2025Updated 10 months ago
- A Framework for Hardware-Aware LLM Exploration☆40Updated this week
- ☆155Jun 13, 2026Updated 2 months ago
- A program for selecting music randomly in DJMAX RESPECT V☆11Jan 30, 2025Updated last year
- A universal adapter including zero-copy Python bindings for Philip Turner's metal flash attention library.☆29Aug 12, 2026Updated 2 weeks ago
- Python package to download FFmpeg binaries from distro servers☆17Aug 11, 2026Updated 2 weeks ago
- Structured output benchmarks comparing DSPy and BAML with different LLMs☆28Dec 23, 2025Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Convert Korean to Katakana☆13Dec 13, 2023Updated 2 years ago
- VALL-E 한국어 버전☆12Aug 22, 2023Updated 3 years ago
- Multilingual-Speech-Synthesis-Voice-Conversion Using Bark + RVC☆14Apr 19, 2025Updated last year
- Lightweight Korean TTS Model based on FastSpeech2☆15Mar 4, 2026Updated 5 months ago
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆355Aug 22, 2026Updated last week
- 대학교 동아리를 위한 그룹웨어☆11May 28, 2021Updated 5 years ago
- Decoding Raymarine's ARCHIVE.FSH files, Garmin's IMG/ADM archives and the TRK subfiles.☆10Oct 22, 2019Updated 6 years ago