The official code implementation for paper "PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs"
☆29May 24, 2025Updated last year
Alternatives and similar repositories for PM-KVQ
Users that are interested in PM-KVQ are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [WSDM'24 Oral] The official implementation of paper <DeSCo: Towards Generalizable and Scalable Deep Subgraph Counting>☆24Mar 11, 2024Updated 2 years ago
- Algorithm-System Co-design: accurate and efficient 2-bit KV cache quantization for LLM Inference.☆17May 20, 2026Updated 2 months ago
- [ECCV24] MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization☆50Nov 27, 2024Updated last year
- [ICML2025] KVTuner: Sensitivity-Aware Layer-wise Mixed Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference☆29Jan 27, 2026Updated 5 months ago
- ☆25Oct 31, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [COLM 2024] SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models☆24Oct 5, 2024Updated last year
- [NeurIPS'25] The official code implementation for paper "R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Tok…☆94Apr 7, 2026Updated 3 months ago
- Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache com…☆15Sep 29, 2025Updated 9 months ago
- The code repository of "MBQ: Modality-Balanced Quantization for Large Vision-Language Models"☆93Mar 17, 2025Updated last year
- ☆16Dec 9, 2023Updated 2 years ago
- ☆30Oct 2, 2025Updated 9 months ago
- [CoLM'25] The official implementation of the paper <MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression>☆159Jan 14, 2026Updated 6 months ago
- [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"☆76Jan 13, 2026Updated 6 months ago
- ☆35Mar 28, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆28Nov 5, 2021Updated 4 years ago
- Code Repository of Evaluating Quantized Large Language Models☆135Sep 8, 2024Updated last year
- [COLM 2025] DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation; 知乎:https://zhuanlan.zhihu.c…☆30Mar 5, 2025Updated last year
- super-resolution; post-training quantization; model compression☆14Nov 10, 2023Updated 2 years ago
- [EMNLP 2024] RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization☆40Sep 24, 2024Updated last year
- A list of papers, docs, codes about diffusion quantization.This repo collects various quantization methods for the Diffusion Models. Welc…☆20Feb 2, 2026Updated 5 months ago
- EQ-Net [ICCV 2023]☆32Aug 15, 2023Updated 2 years ago
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models☆27Apr 2, 2026Updated 3 months ago
- [DATE'23] The official code for paper <CLAP: Locality Aware and Parallel Triangle Counting with Content Addressable Memory>☆24May 25, 2026Updated last month
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Official implementation of "TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization" (Findings of ACL …☆21Jul 25, 2025Updated 11 months ago
- ☆27Jan 20, 2026Updated 6 months ago
- Nsight Systems In Docker☆21Dec 21, 2023Updated 2 years ago
- [ICCV 2025] QuantCache:Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video Generation☆18Sep 26, 2025Updated 9 months ago
- ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression (DAC'25)☆32Feb 26, 2026Updated 4 months ago
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs☆25Nov 11, 2025Updated 8 months ago
- [ICCV-2023] EMQ: Evolving Training-free Proxies for Automated Mixed Precision Quantization☆29Dec 6, 2023Updated 2 years ago
- ☆25Dec 11, 2021Updated 4 years ago
- ☆17Oct 5, 2025Updated 9 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Code for the AAAI 2024 Oral paper "OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Model…☆72Mar 7, 2024Updated 2 years ago
- [ICML 2025] Official PyTorch implementation of "FlatQuant: Flatness Matters for LLM Quantization"☆223Nov 25, 2025Updated 7 months ago
- (ICML-2025) Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers☆20Aug 13, 2025Updated 11 months ago
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆191Updated this week
- Quantization in the Jagged Loss Landscape of Vision Transformers☆13Oct 22, 2023Updated 2 years ago
- [AIMO2] 2nd place solution☆77May 28, 2025Updated last year
- ☆47Nov 25, 2024Updated last year