Activation-aware Singular Value Decomposition for Compressing Large Language Models
β92Oct 22, 2024Updated last year
Alternatives and similar repositories for ASVD4LLM
Users that are interested in ASVD4LLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025π₯] SVD-LLM & [NAACL 2025π₯] SVD-LLM V2β302Aug 28, 2025Updated 10 months ago
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. π The official implementation of https://arxβ¦β31Feb 17, 2025Updated last year
- [ICLR 2025] Palu: Compressing KV-Cache with Low-Rank Projectionβ158Feb 20, 2025Updated last year
- PyTorch code for our paper "AdaSVD: Adaptive Singular Value Decomposition for Large Language Models"β15Mar 9, 2025Updated last year
- [ICML 2025] SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Modelsβ62Aug 9, 2024Updated last year
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- β64Oct 17, 2023Updated 2 years ago
- β129Jan 22, 2024Updated 2 years ago
- [NeurIPS 2024] VeLoRA : Memory Efficient Training using Rank-1 Sub-Token Projectionsβ22Oct 15, 2024Updated last year
- [AAAI 2026] Official implementation of "FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models". If you find this reposiβ¦β17May 1, 2026Updated 2 months ago
- [ICML 2025] Official PyTorch implementation of "FlatQuant: Flatness Matters for LLM Quantization"β223Nov 25, 2025Updated 8 months ago
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Modelsβ31Aug 5, 2025Updated 11 months ago
- β15Nov 7, 2024Updated last year
- Official Implementation of "GRIFFIN: Effective Token Alignment for Faster Speculative Decoding"[NeurIPS 2025]β19May 12, 2025Updated last year
- [ICCV 2025] QuEST: Efficient Finetuning for Low-bit Diffusion Modelsβ60Jun 26, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR'25] R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inferenceβ21Apr 28, 2025Updated last year
- Github Repo for OATS: Outlier-Aware Pruning through Sparse and Low Rank Decompositionβ20Apr 16, 2025Updated last year
- APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decodingβ14Jul 22, 2024Updated 2 years ago
- Official Pytorch Implementation of "Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity"β82Jul 7, 2025Updated last year
- For releasing code related to compression methods for transformers, accompanying our publicationsβ461Jan 16, 2025Updated last year
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantizationβ176Nov 26, 2025Updated 8 months ago
- Code for "Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes"β32Mar 28, 2024Updated 2 years ago
- [EMNLP 2025] AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Modelsβ16Apr 29, 2026Updated 2 months ago
- Official code release for Delta Activations: A Representation for Finetuned Large Language Modelsβ20Sep 5, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β30Jul 22, 2024Updated 2 years ago
- [ICLR 2025] Dobi-SVD : Differentiable SVD for LLM Compression and Some New Perspectives"β54Oct 19, 2025Updated 9 months ago
- [EMNLP 2024] RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantizationβ41Sep 24, 2024Updated last year
- [ICLR 2024 Spotlight] This is the official PyTorch implementation of "EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diβ¦β73Jun 4, 2024Updated 2 years ago
- AFPQ code implementationβ23Nov 6, 2023Updated 2 years ago
- Official implementation of ICML'24 paper "LQER: Low-Rank Quantization Error Reconstruction for LLMs"β19Jul 11, 2024Updated 2 years ago
- Code implementation of GPTAQ (https://arxiv.org/abs/2504.02692)β92Jul 28, 2025Updated 11 months ago
- Code repo for the paper "SpinQuant LLM quantization with learned rotations"β417Feb 14, 2025Updated last year
- Reorder-based post-training quantization for large language modelβ199May 17, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official Repo for SparseLLM: Global Pruning of LLMs (NeurIPS 2024)β70Mar 27, 2025Updated last year
- Analyze the inference of Large Language Models (LLMs). Analyze aspects like computation, storage, transmission, and hardware roofline modβ¦β663Sep 11, 2024Updated last year
- [COLM 2025] Official PyTorch implementation of "Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models"β77Jul 8, 2025Updated last year
- Awesome list for LLM pruning.β297Oct 11, 2025Updated 9 months ago
- [ICML 2024] BiLLM: Pushing the Limit of Post-Training Quantization for LLMsβ235Jan 11, 2025Updated last year
- Benchmark tests supporting the TiledCUDA library.β19Nov 19, 2024Updated last year
- This repository contains code for the MicroAdam paper.β21Dec 14, 2024Updated last year