SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression
☆202May 15, 2026Updated 3 months ago
Alternatives and similar repositories for spectralquant
Users that are interested in spectralquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Sequential Monte Carlo Speculative Decoding☆52Aug 25, 2026Updated 2 weeks ago
- turboquant-based compression engine for LLM KV cache☆61Apr 3, 2026Updated 5 months ago
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆538Updated this week
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,046Apr 23, 2026Updated 4 months ago
- NanoGPT speedrun in JAX. Originally at https://nor-git.pages.dev/modded-nanogpt-jax/☆17Aug 28, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆847Aug 4, 2026Updated last month
- world-model-mcp is a signed audit memory server for AI coding agents, delivered as an MCP tool. Every event Ed25519-signed and Merkle-cha…☆22Aug 27, 2026Updated last week
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Jul 30, 2026Updated last month
- ☆308Apr 5, 2026Updated 5 months ago
- Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)☆88May 29, 2026Updated 3 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,839Updated this week
- TokenSpeed is a speed-of-light LLM inference engine.☆2,103Updated this week
- ☆13Jun 29, 2024Updated 2 years ago
- Algorithms for latent compaction☆265Apr 22, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆27Jul 13, 2026Updated last month
- mKernel: fast multi-node, multi-GPU fused kernels☆270Updated this week
- Delta Attention Residuals - supplementary code and pretrained models☆43May 20, 2026Updated 3 months ago
- ☆45Mar 9, 2026Updated 5 months ago
- Experimental llama.cpp fork for inference research and development☆824Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆339Aug 19, 2026Updated 2 weeks ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 5 months ago
- ☆21Jun 12, 2026Updated 2 months ago
- Jax-based non-linear reconciliation and learning☆17Jun 25, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- a LLM inference engine to run on consumer hardware☆48Apr 15, 2026Updated 4 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 2 months ago
- SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB G…☆210Sep 1, 2026Updated last week
- A demo showcasing Vapi AI into Nextjs apps. A car rental website built with Next.js (App Router), Vercel Postgres, Tailwind CSS, Shadcn U…☆13Jul 14, 2024Updated 2 years ago
- Triton‑style kernel toolkit for MLX plus a small upstream incubator: prototype, benchmark, and upstream fusions for Apple Silicon☆53Mar 31, 2026Updated 5 months ago
- The Geometric OptimizAtion Libraries☆20Mar 12, 2025Updated last year
- ☆147Updated this week
- ☆19Aug 23, 2025Updated last year
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,550Mar 19, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆35Nov 11, 2025Updated 9 months ago
- SQL-like query language and CLI for Qdrant vector search engine☆47Jun 13, 2026Updated 2 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,056Aug 18, 2026Updated 3 weeks ago
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,244Sep 1, 2026Updated last week
- Scripts for managing Debian and RPM package repositories☆17Jan 14, 2026Updated 7 months ago
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.☆19Jun 12, 2026Updated 2 months ago
- ☆17Jun 15, 2026Updated 2 months ago