☆308Apr 5, 2026Updated 5 months ago
Alternatives and similar repositories for turboquant-gpu
Users that are interested in turboquant-gpu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,046Apr 23, 2026Updated 4 months ago
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Experimental llama.cpp fork for inference research and development☆824Updated this week
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆847Aug 4, 2026Updated last month
- turboquant-based compression engine for LLM KV cache☆61Apr 3, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆202May 15, 2026Updated 3 months ago
- Run Qwen3.5-35B-A3B with 1 million token context on a single NVIDIA L4☆29Apr 9, 2026Updated 4 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,056Aug 18, 2026Updated 3 weeks ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,839Updated this week
- The first distributed AGI system. Thousands of autonomous AI agents collaboratively train models, share experiments via P2P gossip, and p…☆2,040Updated this week
- Examples, end-2-end tutorials and apps built using Liquid AI Foundational Models (LFM) and the LEAP SDK☆2,436Updated this week
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆61Jul 30, 2026Updated last month
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆47Aug 17, 2026Updated 3 weeks ago
- ☆15Apr 24, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Agent memory infrastructure: provenance, rollback, lifecycle/supersession, three-layer model (working memory + session archive + wiki). M…☆292May 8, 2026Updated 4 months ago
- LLM inference in C/C++☆2,356Updated this week
- A lightweight inference engine supporting speculative speculative decoding (SSD).☆995May 10, 2026Updated 3 months ago
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆11,690Updated this week
- MiMo-V2.5 NVFP4 4-bit weights + NVFP4 4-bit KV cache + DFlash speculative decoding on 2× NVIDIA DGX Spark — 1M context, 3.4M-token KV poo…☆18Jul 13, 2026Updated last month
- A vector index built on TurboQuant, written in Rust with Python bindings☆16,681Aug 21, 2026Updated 2 weeks ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆19Aug 6, 2026Updated last month
- FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intellige…☆111Aug 18, 2026Updated 3 weeks ago
- An experiential learning fork of TrueAGI's "minecraft-demo"☆17Aug 21, 2026Updated 2 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Security scanning for AI Agents☆33Mar 28, 2026Updated 5 months ago
- Meet IFR: a bio-inspired engine solving RAG’s biggest flaws. It achieves true O(1) scaling latency stays <5ms even as data grows 1000x. W…☆15Apr 3, 2026Updated 5 months ago
- ☆28May 10, 2026Updated 3 months ago
- Autonomous FPV drone control with pure vision input☆17Dec 5, 2025Updated 9 months ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆26Aug 22, 2026Updated 2 weeks ago
- Pi extension that enables agents to look things up via natural language query.☆32May 28, 2026Updated 3 months ago
- 🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.☆2,202Updated this week
- [MLsys2026 Best Paper]: https://arxiv.org/abs/2506.08276. RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, a…☆12,895Updated this week
- A simple library for generating instruction tuning datasets locally☆108Aug 21, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆7,021Jul 20, 2026Updated last month
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆74Aug 2, 2026Updated last month
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,763Updated this week
- ☆22Oct 14, 2024Updated last year
- 어느 고등학생의 심플한 확률론적 앵무새 만들기☆19Aug 3, 2026Updated last month
- Runs 405B LLMs on 8GB VRAM☆3,059Updated this week
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆1,273Updated this week