LLM KV Cache compression - K+V dual compression, 73-99% VRAM savings, zero accuracy loss
☆57Mar 30, 2026Updated 4 months ago
Alternatives and similar repositories for polarquant-kv
Users that are interested in polarquant-kv are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☕️ A vscode extension for netron, support *.pdmodel, *.nb, *.onnx, *.pb, *.h5, *.tflite, *.pth, *.pt, *.mnn, *.param, etc.☆14Jun 4, 2023Updated 3 years ago
- A Rust library for data structures living in shared memory.☆14Jun 25, 2020Updated 6 years ago
- ☆21Mar 22, 2021Updated 5 years ago
- [HPCA 2026] A GPU-optimized system for efficient long-context LLMs decoding with low-bit KV cache.☆96May 14, 2026Updated 2 months ago
- ☆49Apr 15, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 中科大郑启龙2021年并行程序设计课程实验☆11Jan 15, 2022Updated 4 years ago
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated 11 months ago
- Super fast accurate face detector ! SCRFD(CVPR 2021) with MNN/TNN/NCNN/ONNXRuntime C++.☆21Jan 12, 2022Updated 4 years ago
- Cute layout visualization☆45Jan 18, 2026Updated 6 months ago
- a simple API to use CUPTI☆10Aug 19, 2025Updated 11 months ago
- Distilling Knowledge via Intermediate Classifiers☆16Oct 3, 2021Updated 4 years ago
- PyTorch implementation of Language model compression with weighted low-rank factorization☆14Jun 28, 2023Updated 3 years ago
- ☆18Sep 23, 2025Updated 10 months ago
- LLM Inference FrameWork☆32May 20, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆24Apr 10, 2026Updated 4 months ago
- Fully open reproduction of DeepSeek-R1☆11Mar 24, 2025Updated last year
- 集中管理所有的prompt。☆14Nov 27, 2024Updated last year
- 🖖 图谱式笔记系统,旨在提高个人笔记的使用率!☆11Jan 17, 2021Updated 5 years ago
- ☆21Apr 27, 2026Updated 3 months ago
- ☆32Jul 2, 2025Updated last year
- Expert Specialization MoE Solution based on CUTLASS☆27Apr 14, 2026Updated 3 months ago
- Official implementation for LaCo (EMNLP 2024 Findings)☆22Oct 3, 2024Updated last year
- A skill for automatically optimizing CUDA code.☆42Mar 26, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- VSS: A Storage System for Video Analytics☆13Jul 9, 2021Updated 5 years ago
- [AAAI-25] Official repository of "Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object De…☆21Dec 27, 2024Updated last year
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- KFunca: A minimalist, high-performance GPU-based automatic differentiation framework☆31Aug 14, 2025Updated 11 months ago
- An arbitrary-precision integer and decimal library for Mojo, also with a 128-bit fixed-point decimal type.☆26Updated this week
- Some funny cute/cuteDSL code snippets☆33Mar 2, 2026Updated 5 months ago
- A tutorial for CUDA&PyTorch☆490Mar 23, 2026Updated 4 months ago
- 🍎 One kernel a day keeps high latency away. A hands-on CUDA learning path featuring a rich collection of kernels, from the basics to pea…☆207Updated this week
- my solution for UC Berkeley AI projects pacman☆11Jul 25, 2020Updated 6 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A rust-version of NVIDIA BlueField DOCA kit.☆13Jun 11, 2023Updated 3 years ago
- Mojo Miji | A guide to Mojo programming language from a Pythonista's perspective | Mojo 秘籍☆34Jun 14, 2026Updated last month
- Official Codebase for "Aligning Diffusion Behaviors with Q-functions for Efficient Continuous Control" (NeurIPS 2024)☆15Oct 29, 2024Updated last year
- Inference Llama 2 in one file of pure Cuda☆17Aug 20, 2023Updated 2 years ago
- 分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等☆3,423Updated this week
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- 计算机网络课程设计, 基于TCP协议的简易聊天机器人, 开发语言Python3, 初期版本只能在终端中运行(CLI), 最终完成版为客户端编写了"简陋"的图形界面, 使用Qt5(即PyQt5)实现☆10Jun 17, 2019Updated 7 years ago