☆269Apr 5, 2026Updated 4 months ago
Alternatives and similar repositories for turboquant-gpu
Users that are interested in turboquant-gpu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,045Apr 23, 2026Updated 3 months ago
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Experimental llama.cpp fork for inference research and development☆759Updated this week
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆842Aug 4, 2026Updated 2 weeks ago
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆203May 15, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Run Qwen3.5-35B-A3B with 1 million token context on a single NVIDIA L4☆28Apr 9, 2026Updated 4 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,643May 10, 2026Updated 3 months ago
- LLM speculative inference server for consumer & heterogeneous hardware☆2,762Updated this week
- Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.☆44Updated this week
- The first distributed AGI system. Thousands of autonomous AI agents collaboratively train models, share experiments via P2P gossip, and p…☆2,024Updated this week
- Examples, end-2-end tutorials and apps built using Liquid AI Foundational Models (LFM) and the LEAP SDK☆2,404Updated this week
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆60Jul 30, 2026Updated 2 weeks ago
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆40Updated this week
- ☆15Apr 24, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Agent memory infrastructure: provenance, rollback, lifecycle/supersession, three-layer model (working memory + session archive + wiki). M…☆292May 8, 2026Updated 3 months ago
- LLM inference in C/C++☆2,291Updated this week
- A lightweight inference engine supporting speculative speculative decoding (SSD).☆989May 10, 2026Updated 3 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,199Updated this week
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated last month
- A vector index built on TurboQuant, written in Rust with Python bindings☆14,853Updated this week
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆18Aug 6, 2026Updated last week
- FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intellige…☆110Aug 3, 2026Updated 2 weeks ago
- An experiential learning fork of TrueAGI's "minecraft-demo"☆17Oct 21, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Security scanning for AI Agents☆33Mar 28, 2026Updated 4 months ago
- Meet IFR: a bio-inspired engine solving RAG’s biggest flaws. It achieves true O(1) scaling latency stays <5ms even as data grows 1000x. W…☆15Apr 3, 2026Updated 4 months ago
- ☆28May 10, 2026Updated 3 months ago
- Lightweight tools for quick and easy LLM demo's☆28Sep 22, 2024Updated last year
- 🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.☆2,164Updated this week
- [MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on …☆12,789Jul 31, 2026Updated 2 weeks ago
- A simple library for generating instruction tuning datasets locally☆95Jun 10, 2026Updated 2 months ago
- ☆7,004Jul 20, 2026Updated 3 weeks ago
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆73Aug 2, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,722Mar 27, 2026Updated 4 months ago
- ☆22Oct 14, 2024Updated last year
- Runs 405B LLMs on 8GB VRAM☆3,056Apr 2, 2026Updated 4 months ago
- Self-hosted multi-agent coordination using Matrix☆18Feb 5, 2026Updated 6 months ago
- The cli of dataverse☆118Mar 22, 2026Updated 4 months ago
- NebulaFlow is for visually designing and running developer workflows as node graphs (CLI, LLM, control‑flow, previews). Build and execute…☆26Jul 19, 2026Updated 3 weeks ago
- Knowledgeable Embedding: Injecting dynamically updatable entity knowledge into embeddings to enhance RAG☆15Aug 31, 2025Updated 11 months ago