☆259Apr 5, 2026Updated 3 months ago
Alternatives and similar repositories for turboquant-gpu
Users that are interested in turboquant-gpu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,041Apr 23, 2026Updated 3 months ago
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 4 months ago
- Experimental llama.cpp fork for inference research and development☆721Updated this week
- turboquant-based compression engine for LLM KV cache☆62Apr 3, 2026Updated 3 months ago
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆197May 15, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Run Qwen3.5-35B-A3B with 1 million token context on a single NVIDIA L4☆28Apr 9, 2026Updated 3 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,547May 10, 2026Updated 2 months ago
- Fast LLM speculative inference server for consumer hardware.☆2,694Updated this week
- Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.☆40May 26, 2026Updated 2 months ago
- The first distributed AGI system. Thousands of autonomous AI agents collaboratively train models, share experiments via P2P gossip, and p…☆2,013Updated this week
- Examples, end-2-end tutorials and apps built using Liquid AI Foundational Models (LFM) and the LEAP SDK☆2,143Updated this week
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆23Jul 12, 2026Updated 2 weeks ago
- LLM inference in C/C++☆2,203Jul 22, 2026Updated last week
- A lightweight inference engine supporting speculative speculative decoding (SSD).☆977May 10, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LMCache: Supercharge Your LLM with the Fastest KV Cache Layer☆10,938Updated this week
- A vector index built on TurboQuant, written in Rust with Python bindings☆14,494Updated this week
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆17Updated this week
- FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intellige…☆109Jun 3, 2026Updated last month
- An experiential learning fork of TrueAGI's "minecraft-demo"☆17Oct 21, 2024Updated last year
- Security scanning for AI Agents☆33Mar 28, 2026Updated 4 months ago
- Meet IFR: a bio-inspired engine solving RAG’s biggest flaws. It achieves true O(1) scaling latency stays <5ms even as data grows 1000x. W…☆15Apr 3, 2026Updated 3 months ago
- ☆28May 10, 2026Updated 2 months ago
- Lightweight tools for quick and easy LLM demo's☆28Sep 22, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU pl…☆24Updated this week
- Pi extension that enables agents to look things up via natural language query.☆32May 28, 2026Updated 2 months ago
- 🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.☆2,132Updated this week
- [MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on …☆12,743Updated this week
- A simple library for generating instruction tuning datasets locally☆90Jun 10, 2026Updated last month
- ☆7,004Jul 20, 2026Updated last week
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆71Updated this week
- TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration☆1,702Mar 27, 2026Updated 4 months ago
- 어느 고등학생의 심플한 확률론적 앵무새 만들기☆19Sep 2, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Runs 405B LLMs on 8GB VRAM☆3,043Apr 2, 2026Updated 3 months ago
- extract all your personal data history from cursor, codex, claude-code, windsurf, and trae☆843Jan 21, 2026Updated 6 months ago
- The cli of dataverse☆118Mar 22, 2026Updated 4 months ago
- NebulaFlow is for visually designing and running developer workflows as node graphs (CLI, LLM, control‑flow, previews). Build and execute…☆26Jul 19, 2026Updated last week
- Knowledgeable Embedding: Injecting dynamically updatable entity knowledge into embeddings to enhance RAG☆15Aug 31, 2025Updated 10 months ago
- OmniDaemon is a Universal Event-Driven Runtime for AI Agents, it's framework-agnostic, event-driven runtime that turns AI agents into pr…☆55Dec 15, 2025Updated 7 months ago
- ☆50Apr 12, 2026Updated 3 months ago