TurboQuant 3-bit KV-cache quantization for llama.cpp
☆56Jul 30, 2026Updated last month
Alternatives and similar repositories for llama-turboquant
Users that are interested in llama-turboquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AutoMapper website☆14Aug 19, 2020Updated 6 years ago
- Generate images locally with Bonsai 1-bit and ternary image generation models.☆29Jun 12, 2026Updated 3 months ago
- Recording models☆13Sep 19, 2023Updated 2 years ago
- convert pytorch model to ncnn☆13Dec 5, 2018Updated 7 years ago
- BIQU thumbnail generator python script for the BTT series of TFT's. The script can be used directly in PrusaSlicer or on it's own.This sc…☆12Feb 3, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A collection of container images used in CI across various opencontainers projects☆18Mar 22, 2023Updated 3 years ago
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆25Apr 1, 2025Updated last year
- Proxmox GPU Passthrough Guide for AMD Ryzen AI Max+ 395 (Strix Halo / 8060S)☆50Sep 12, 2025Updated last year
- Unique Entware packages written in GO☆24Updated this week
- TurboQuant llama.cpp fork with optimized turbo4 kernels for Gemma 4 D=256/512 heads — lazy K/V, batch decode, warp-cooperative write. 120…☆36Apr 5, 2026Updated 5 months ago
- ☆15Jun 6, 2025Updated last year
- LLM inference in C/C++☆2,374Updated this week
- The fastest way to run Qwen3.8 27B on Strix Halo (gfx1151)☆60Sep 5, 2026Updated last week
- ☆31Aug 28, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Jul 21, 2023Updated 3 years ago
- This project is intended to build and deploy an SNPE model on Qualcomm Devices, which are having unsupported layers which are not part of…☆10Oct 4, 2021Updated 4 years ago
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,047Apr 23, 2026Updated 4 months ago
- LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model: sy…☆31Updated this week
- Converts CLIP models to ONNX☆11Jan 17, 2023Updated 3 years ago
- minimal C implementation of speculative decoding based on llama2.c☆30Jul 15, 2024Updated 2 years ago
- Local-first AI citizens on your own hardware. Teams of continuously-learning personas (Rust core, llama.cpp fork, LoRA genomes) that coll…☆29Updated this week
- Write data migration logic in code so you can change the shape of your data confidently as your app evolves☆15Sep 29, 2023Updated 2 years ago
- ☆73Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 基于.net core2.1 搭建的框架的中小型项目框架,sql server数据库,同时支持,mysql数据库,,,需要自己稍微扩展下,,,可以擦模考sql server的实现 迁移文件已经生成,可以直接运行即可跑起来. 结合docker发布可以参考这里: https…☆13Dec 4, 2018Updated 7 years ago
- LDC: Lightweight Dense CNN for Edge DetectionのPythonでのONNX推論サンプル☆15May 6, 2023Updated 3 years ago
- 开源网址导航☆13Apr 15, 2021Updated 5 years ago
- With a few words and a click of a button, quickly get an engaging, high quality video. (And optionally save and share it!)☆18May 4, 2025Updated last year
- ⚡ Sao chép Google Drive siêu tốc 🚀 Xử lý đa luồng giúp tăng tốc độ sao chép dữ liệu 🔄 Tự động khôi phục và tiếp tục công việc khi gặp …☆20Jul 20, 2026Updated last month
- nack keyboard☆11Mar 15, 2024Updated 2 years ago
- Fork of RecurrentGPT with modifications☆10Sep 18, 2024Updated last year
- Samples on how to use filters to authenticate ASP .Net Web Api 2☆10Oct 4, 2016Updated 9 years ago
- RAG-Fusion implementation using Langchain, Weaviate and OpenAI☆13Oct 31, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- RedisMQDemo☆11Jun 13, 2019Updated 7 years ago
- Python implementation for regime-dependent portfolio optimization☆17Oct 14, 2023Updated 2 years ago
- The .NET library to integrate Google's EmbeddingGemma-300m model into .NET projects☆15Apr 4, 2026Updated 5 months ago
- ffmpeg+cuvid+tensorrt+multicamera☆12Dec 31, 2024Updated last year
- Copy the tick and history from the MetaTrade 4 to MetaTrader 5☆19Jul 1, 2020Updated 6 years ago
- Distribute and run LLMs with a single file.☆25May 13, 2025Updated last year
- This app uses OpenAI's LLM model to answer questions about your PDF file. Upload your PDF file and ask questions about it. The app will r…☆13May 13, 2025Updated last year