turboquant-based compression engine for LLM KV cache
☆62Apr 3, 2026Updated 3 months ago
Alternatives and similar repositories for turboquant_cutile
Users that are interested in turboquant_cutile are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Repository for GPU related kernels for learning/testing purposes☆19May 27, 2026Updated 2 months ago
- ☆19Apr 26, 2026Updated 3 months ago
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆17Jul 11, 2026Updated 2 weeks ago
- Real-Time Mock Technical Interview Platform☆11Sep 2, 2025Updated 10 months ago
- A codebase for pretraining multi-billion-scale sparse GPTs.☆24Feb 9, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆197May 15, 2026Updated 2 months ago
- The official code for "Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation" | [MM2…☆14Dec 7, 2024Updated last year
- A Triton-only attention backend for vLLM☆27Jul 14, 2026Updated 2 weeks ago
- ☆20Mar 17, 2026Updated 4 months ago
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆21Nov 28, 2025Updated 8 months ago
- a Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization in pure C.☆24Jul 6, 2024Updated 2 years ago
- Write a fast kernel and see how you compare against the best humans and AI on gpumode.com☆105Updated this week
- AirLLM 70B inference with single 4GB GPU☆21Jun 27, 2025Updated last year
- ☆35Jun 28, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆35May 22, 2026Updated 2 months ago
- 1bit llama.cpp gguf weights paired with turboquant 4 bit kv cache☆23Apr 4, 2026Updated 3 months ago
- ☆259Apr 5, 2026Updated 3 months ago
- Efficient Long-context Language Model Training by Core Attention Disaggregation☆106Apr 7, 2026Updated 3 months ago
- Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free…☆25Updated this week
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆123Apr 17, 2026Updated 3 months ago
- ☆390Apr 16, 2026Updated 3 months ago
- Code for the paper *Attention Drift: What Speculative Decoding Models Learn*.☆28May 12, 2026Updated 2 months ago
- GPU kernel benchmarking☆47Jun 10, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 一个普通的网站☆17May 18, 2025Updated last year
- a mini TPU with floating point arithmetic☆53Dec 22, 2025Updated 7 months ago
- A dynamic binary instrumentation tool for tracing and analyzing CUDA kernel instructions.☆76Updated this week
- DeeperGEMM: crazy optimized version☆86May 5, 2025Updated last year
- Github mirror of trition-lang/triton repo.☆181Updated this week
- [ICML2026] Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization☆60Updated this week
- This repo contains Lyra AI's work in the E-Commerce Hackathon organized by Trendyol and Teknofest.☆11Nov 6, 2024Updated last year
- Distributed multi-agent framework for event-driven, graph-based computation. Elixir/Python, NATS event streaming, modular operator/XCS ar…☆14Mar 25, 2026Updated 4 months ago
- ☆37Aug 7, 2025Updated 11 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- One command · One Microsoft login · Zero repeated auth Hours of uninterrupted access to NYU Torch from your terminal and IDE.☆15Jul 17, 2026Updated last week
- A graph database library for iOS and MacOS.☆14Mar 4, 2018Updated 8 years ago
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆369Jul 9, 2026Updated 3 weeks ago
- ☆55Updated this week
- ☆18Dec 3, 2025Updated 7 months ago
- ☆91Oct 17, 2025Updated 9 months ago
- 机器人人工智能,优达学城cs373作业。 Artificial Intelligence for Robotics, this repository contains all the homework…☆12Nov 12, 2017Updated 8 years ago