SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression
☆203May 15, 2026Updated 3 months ago
Alternatives and similar repositories for spectralquant
Users that are interested in spectralquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Sequential Monte Carlo Speculative Decoding☆52Aug 10, 2026Updated last week
- turboquant-based compression engine for LLM KV cache☆62Apr 3, 2026Updated 4 months ago
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆523Updated this week
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,045Apr 23, 2026Updated 3 months ago
- Triton kernels for dynamic causal short convolutions.☆27Jun 4, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- NanoGPT speedrun in JAX. Originally at https://nor-git.pages.dev/modded-nanogpt-jax/☆17Aug 28, 2025Updated 11 months ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆842Aug 4, 2026Updated 2 weeks ago
- world-model-mcp is a signed audit memory server for AI coding agents, delivered as an MCP tool. Every event Ed25519-signed and Merkle-cha…☆22Aug 2, 2026Updated 2 weeks ago
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Jul 30, 2026Updated 2 weeks ago
- Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)☆88May 29, 2026Updated 2 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆1,926Updated this week
- LLM speculative inference server for consumer & heterogeneous hardware☆2,762Updated this week
- Algorithms for latent compaction☆262Apr 22, 2026Updated 3 months ago
- ☆27Jul 13, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- mKernel: fast multi-node, multi-GPU fused kernels☆266Updated this week
- ☆16Updated this week
- ☆45Mar 9, 2026Updated 5 months ago
- Experimental llama.cpp fork for inference research and development☆759Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆328Jul 1, 2026Updated last month
- ☆21Jun 12, 2026Updated 2 months ago
- Jax-based non-linear reconciliation and learning☆17Jun 25, 2026Updated last month
- a LLM inference engine to run on consumer hardware☆47Apr 15, 2026Updated 4 months ago
- SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB G…☆207Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆37Jun 12, 2026Updated 2 months ago
- Implementation for Revela: Dense Retriever Learning via Language Modeling - ICLR 2026 Oral☆20Mar 26, 2026Updated 4 months ago
- Triton‑style kernel toolkit for MLX plus a small upstream incubator: prototype, benchmark, and upstream fusions for Apple Silicon☆50Mar 31, 2026Updated 4 months ago
- The Geometric OptimizAtion Libraries☆20Mar 12, 2025Updated last year
- ☆145Mar 9, 2026Updated 5 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,529Mar 19, 2026Updated 5 months ago
- ☆19Aug 23, 2025Updated 11 months ago
- ☆35Nov 11, 2025Updated 9 months ago
- SQL-like query language and CLI for Qdrant vector search engine☆46Jun 13, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Scripts for managing Debian and RPM package repositories☆17Jan 14, 2026Updated 7 months ago
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.☆19Jun 12, 2026Updated 2 months ago
- ☆17Jun 15, 2026Updated 2 months ago
- Code for the paper *Attention Drift: What Speculative Decoding Models Learn*.☆28May 12, 2026Updated 3 months ago
- Using modal.com to process FineWeb-edu data☆20Apr 11, 2026Updated 4 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆5,643May 10, 2026Updated 3 months ago
- ☆86Updated this week