☆21Feb 5, 2024Updated 2 years ago
Alternatives and similar repositories for SqueezeLLM-gradients
Users that are interested in SqueezeLLM-gradients are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2024] SqueezeLLM: Dense-and-Sparse Quantization☆724Aug 13, 2024Updated 2 years ago
- [ICLR25] STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs☆20Jun 3, 2025Updated last year
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆25Dec 5, 2025Updated 8 months ago
- (ICML-2025) Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers☆21Aug 13, 2025Updated last year
- [NeurIPS 2024] KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization☆432Aug 13, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models☆23Mar 15, 2024Updated 2 years ago
- ☆43Jan 30, 2024Updated 2 years ago
- [ICLR 2026] This is the official PyTorch implementation of "QVGen: Pushing the Limit of Quantized Video Generative Models".☆32Feb 11, 2026Updated 6 months ago
- ☆11May 24, 2024Updated 2 years ago
- [COLM 2025] DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation; 知乎:https://zhuanlan.zhihu.c…☆30Mar 5, 2025Updated last year
- ☆82Jul 21, 2022Updated 4 years ago
- PyTorch implementation of Language model compression with weighted low-rank factorization☆14Jun 28, 2023Updated 3 years ago
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- Artifact evaluation for HPCA'24 paper Lightening-Transformer: A Dynamically-operated Optically-interconnected Photonic Transformer Accele…☆11Mar 3, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆19Jan 3, 2025Updated last year
- ☆11Nov 14, 2023Updated 2 years ago
- Piecewise-Affine Regularized Quantization☆19Feb 5, 2026Updated 6 months ago
- ☆45Nov 1, 2022Updated 3 years ago
- Supporting code for "LLMs for your iPhone: Whole-Tensor 4 Bit Quantization"☆11Mar 31, 2024Updated 2 years ago
- [ACL 2025] Squeezed Attention: Accelerating Long Prompt LLM Inference☆58Nov 20, 2024Updated last year
- Code for paper: "QuIP: 2-Bit Quantization of Large Language Models With Guarantees"☆402Feb 24, 2024Updated 2 years ago
- Open Source Projects from Pallas Lab☆21Oct 10, 2021Updated 4 years ago
- [ICLR'25] Code for KaSA, an official implementation of "KaSA: Knowledge-Aware Singular-Value Adaptation of Large Language Models"☆22Jan 16, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆11Apr 5, 2023Updated 3 years ago
- Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 30+ benchmarks☆15Feb 17, 2025Updated last year
- O'Reilly Course, In-Memory Computing Essentials☆10Oct 16, 2020Updated 5 years ago
- ☆19Feb 4, 2025Updated last year
- This repo contains the source code for: Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs☆44Aug 14, 2024Updated 2 years ago
- This repository contains low-bit quantization papers from 2020 to 2026 on top conference.☆206Jun 25, 2026Updated last month
- Code for paper: "QuIP: 2-Bit Quantization of Large Language Models With Guarantees" adapted for Llama models☆40Aug 4, 2023Updated 3 years ago
- ☆14Oct 3, 2024Updated last year
- Awesome LLM compression research papers and tools.☆1,862Jun 30, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆53May 13, 2024Updated 2 years ago
- [NeurIPS 2024 Oral🔥] DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs.☆187Apr 24, 2026Updated 3 months ago
- [TrimKV] Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs - [DBTrimKV] Make Each Token Count: Towards Improving Lo…☆18Jul 26, 2026Updated 3 weeks ago
- Code Repository of Evaluating Quantized Large Language Models☆135Sep 8, 2024Updated last year
- ☆19Dec 10, 2021Updated 4 years ago
- [ICML 2024 Oral] Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs☆132Jul 4, 2025Updated last year
- 😎 Awesome papers on token redundancy reduction☆14Mar 12, 2025Updated last year