SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression
☆197May 15, 2026Updated 2 months ago
Alternatives and similar repositories for spectralquant
Users that are interested in spectralquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Public benchmark results from Kernel Arena, a leaderboard for LLM-generated AI accelerator kernels.☆20Mar 11, 2026Updated 4 months ago
- turboquant-based compression engine for LLM KV cache☆62Apr 3, 2026Updated 3 months ago
- SutroYaro — Sutro Group research workspace for energy-efficient AI training. Point any coding agent at the repo and it becomes a research…☆15May 29, 2026Updated last month
- REAM: Merging Improves Pruning of Experts in LLMs☆21Apr 16, 2026Updated 3 months ago
- A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official …☆424Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,037Apr 23, 2026Updated 2 months ago
- Triton kernels for dynamic causal short convolutions.☆24Jun 4, 2026Updated last month
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆826Jul 14, 2026Updated last week
- Temporal knowledge graph for AI coding agents. 28 MCP tools, 10 runtime adapters (Claude Code, Cursor, Codex, Hermes, Continue, OpenClaw,…☆19Jul 11, 2026Updated last week
- Official repository for Parallax (Parameterized Local Linear Attention)☆65Jul 7, 2026Updated 2 weeks ago
- ☆259Apr 5, 2026Updated 3 months ago
- Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)☆86May 29, 2026Updated last month
- TokenSpeed is a speed-of-light LLM inference engine.☆1,638Updated this week
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆13Jun 29, 2024Updated 2 years ago
- Algorithms for latent compaction☆257Apr 22, 2026Updated 2 months ago
- ☆22Jul 13, 2026Updated last week
- mKernel: fast multi-node, multi-GPU fused kernels☆251Jun 21, 2026Updated last month
- Delta Attention Residuals - supplementary code and pretrained models☆40May 20, 2026Updated 2 months ago
- ☆15May 20, 2026Updated 2 months ago
- ☆45Mar 9, 2026Updated 4 months ago
- LLAMA Turboquant implementation with CUDA support☆703Updated this week
- [ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference☆325Jul 1, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆22Apr 3, 2026Updated 3 months ago
- ☆21Jun 12, 2026Updated last month
- Jax-based non-linear reconciliation and learning☆17Jun 25, 2026Updated 3 weeks ago
- a LLM inference engine to run on consumer hardware☆46Apr 15, 2026Updated 3 months ago
- SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB G…☆52Updated this week
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆37Jun 12, 2026Updated last month
- Implementation for Revela: Dense Retriever Learning via Language Modeling - ICLR 2026 Oral☆20Mar 26, 2026Updated 3 months ago
- A demo showcasing Vapi AI into Nextjs apps. A car rental website built with Next.js (App Router), Vercel Postgres, Tailwind CSS, Shadcn U…☆13Jul 14, 2024Updated 2 years ago
- Triton‑style kernel toolkit for MLX plus a small upstream incubator: prototype, benchmark, and upstream fusions for Apple Silicon☆47Mar 31, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The Geometric OptimizAtion Libraries☆19Mar 12, 2025Updated last year
- Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free…☆26Updated this week
- ☆35Nov 11, 2025Updated 8 months ago
- ☆19Aug 23, 2025Updated 10 months ago
- SQL-like query language and CLI for Qdrant vector search engine☆46Jun 13, 2026Updated last month
- FlashKDA: high-performance Kimi Delta Attention kernels☆462May 26, 2026Updated last month
- ☆44Jan 9, 2025Updated last year