☆307Apr 5, 2026Updated 5 months ago
Alternatives and similar repositories for turboquant-gpu
Users that are interested in turboquant-gpu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,052Apr 23, 2026Updated 5 months ago
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Sep 18, 2026Updated 2 weeks ago
- Experimental llama.cpp fork for inference research and development☆879Updated this week
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆854Aug 4, 2026Updated last month
- turboquant-based compression engine for LLM KV cache☆61Apr 3, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆203May 15, 2026Updated 4 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,131Aug 18, 2026Updated last month
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,887Updated this week
- Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.☆45Aug 16, 2026Updated last month
- The first distributed AGI system. Thousands of autonomous AI agents collaboratively train models, share experiments via P2P gossip, and p…☆2,066Updated this week
- Examples, end-2-end tutorials and apps built using Liquid AI Foundational Models (LFM) and the LEAP SDK☆2,530Updated this week
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆51Aug 17, 2026Updated last month
- Agent memory infrastructure: provenance, rollback, lifecycle/supersession, three-layer model (working memory + session archive + wiki). M…☆291May 8, 2026Updated 4 months ago
- LLM inference in C/C++☆2,414Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A lightweight inference engine supporting speculative speculative decoding (SSD).☆1,009May 10, 2026Updated 4 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,946Updated this week
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated 2 months ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆20Aug 6, 2026Updated last month
- FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intellige…☆115Aug 18, 2026Updated last month
- A vector index built on TurboQuant, written in Rust with Python bindings☆17,269Updated this week
- An experiential learning fork of TrueAGI's "minecraft-demo"☆17Aug 21, 2026Updated last month
- Security scanning for AI Agents☆33Mar 28, 2026Updated 6 months ago
- ☆28Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Lightweight tools for quick and easy LLM demo's☆28Sep 22, 2024Updated 2 years ago
- Pi extension that enables agents to look things up via natural language query.☆32May 28, 2026Updated 4 months ago
- 🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.☆2,295Updated this week
- A simple library for generating instruction tuning datasets locally☆117Sep 20, 2026Updated last week
- ☆7,031Jul 20, 2026Updated 2 months ago
- Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents☆55Sep 6, 2026Updated 3 weeks ago
- ☆22Oct 14, 2024Updated last year
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆77Aug 2, 2026Updated 2 months ago
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,788Sep 3, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 어느 고등학생의 심플한 확률론적 앵무새 만들기☆19Aug 3, 2026Updated 2 months ago
- Runs 405B LLMs on 8GB VRAM☆3,060Updated this week
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆1,288Sep 26, 2026Updated last week
- The cli of dataverse☆118Aug 20, 2026Updated last month
- NebulaFlow is for visually designing and running developer workflows as node graphs (CLI, LLM, control‑flow, previews). Build and execute…☆26Jul 19, 2026Updated 2 months ago
- Knowledgeable Embedding: Injecting dynamically updatable entity knowledge into embeddings to enhance RAG☆15Aug 31, 2025Updated last year
- OmniDaemon is a Universal Event-Driven Runtime for AI Agents, it's framework-agnostic, event-driven runtime that turns AI agents into pr…☆56Dec 15, 2025Updated 9 months ago