A safetensors extension to efficiently store sparse quantized tensors on disk
☆302Jul 18, 2026Updated this week
Alternatives and similar repositories for compressed-tensors
Users that are interested in compressed-tensors are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM☆3,561Updated this week
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆629Updated this week
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.☆1,109Sep 4, 2024Updated last year
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆191Updated this week
- ☆210May 5, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Fast low-bit matmul kernels in Triton☆477Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆267Dec 4, 2025Updated 7 months ago
- Boosting 4-bit inference kernels with 2:4 Sparsity☆96Sep 4, 2024Updated last year
- DeeperGEMM: crazy optimized version☆86May 5, 2025Updated last year
- FlashInfer: Kernel Library for LLM Serving☆5,983Updated this week
- PyTorch native quantization and sparsity for training and inference☆2,906Updated this week
- Fast Hadamard transform in CUDA, with a PyTorch interface☆338Mar 10, 2026Updated 4 months ago
- ☆114Feb 26, 2026Updated 4 months ago
- A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative…☆3,262Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- Code for Neurips24 paper: QuaRot, an end-to-end 4-bit inference of large language models.☆523Nov 26, 2024Updated last year
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI tr…☆121Updated this week
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆15Jul 11, 2026Updated last week
- Code for data-aware compression of DeepSeek models☆75Dec 11, 2025Updated 7 months ago
- Model Compression Toolbox for Large Language Models and Diffusion Models☆794Aug 14, 2025Updated 11 months ago
- Fast Matrix Multiplications for Lookup Table-Quantized LLMs☆391Apr 13, 2025Updated last year
- extensible collectives library in triton☆97Mar 31, 2025Updated last year
- Triton-based implementation of Sparse Mixture of Experts.☆281Oct 3, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆126Mar 18, 2026Updated 4 months ago
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on H…☆3,434Updated this week
- A pytorch quantization backend for optimum☆1,046Updated this week
- ☆91Jan 23, 2025Updated last year
- AutoAWQ implements the AWQ algorithm for 4-bit quantization with a 2x speedup during inference. Documentation:☆2,350May 11, 2025Updated last year
- ☆170Dec 27, 2024Updated last year
- Accelerating MoE with IO and Tile-aware Optimizations☆731Jul 4, 2026Updated 2 weeks ago
- LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM…☆1,208Updated this week
- mKernel: fast multi-node, multi-GPU fused kernels☆251Jun 21, 2026Updated 3 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆269Jul 11, 2024Updated 2 years ago
- ☆14Jul 13, 2025Updated last year
- ☆135May 29, 2025Updated last year
- [MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Se…☆850Mar 6, 2025Updated last year
- Official implementation of Half-Quadratic Quantization (HQQ)☆949Feb 26, 2026Updated 4 months ago
- NVIDIA Inference Xfer Library (NIXL)☆1,138Updated this week
- ☆114Apr 19, 2024Updated 2 years ago