This repository provides the official implementation of QSVD, a method for efficient low-rank approximation that unifies Query-Key-Value (QKV) weight compression in low-precision Vision-Language Models (VLMs).
☆29May 16, 2026Updated 4 months ago
Alternatives and similar repositories for QSVD
Users that are interested in QSVD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] Dobi-SVD : Differentiable SVD for LLM Compression and Some New Perspectives"☆54Oct 19, 2025Updated 11 months ago
- The code repository of "MBQ: Modality-Balanced Quantization for Large Vision-Language Models"☆96Mar 17, 2025Updated last year
- [AAAI 2026] Official implementation of "FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models". If you find this reposi…☆19Aug 26, 2026Updated 3 weeks ago
- [NeurIPS 2025 (spotlight)] HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs☆16Dec 17, 2025Updated 9 months ago
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization☆44Aug 13, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆19Updated this week
- [ICLR 2026 Oral] Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation☆104May 8, 2026Updated 4 months ago
- TA's implementation for the project of Computer Architecture and Intelligent Chip Design (23 Spring)☆10May 20, 2023Updated 3 years ago
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated last year
- ☆32Feb 6, 2026Updated 7 months ago
- PyTorch code for our paper "AdaSVD: Adaptive Singular Value Decomposition for Large Language Models"☆16Mar 9, 2025Updated last year
- 6 DoF Directional Room Impulse Response (RIR) with Dense Loudspeaker Grid☆17Aug 31, 2023Updated 3 years ago
- Code implementation of GPTAQ (https://arxiv.org/abs/2504.02692)☆96Jul 28, 2025Updated last year
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆40Dec 24, 2025Updated 8 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICRA 2025] Fast Global Localization on Neural Radiance Field☆17Jun 30, 2025Updated last year
- ☆13Aug 17, 2020Updated 6 years ago
- ☆16Oct 28, 2025Updated 10 months ago
- [ICML 2025] Official PyTorch implementation of "FlatQuant: Flatness Matters for LLM Quantization"☆228Nov 25, 2025Updated 9 months ago
- DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails☆34Feb 26, 2025Updated last year
- LLGS: Illuminating Gaussian Splatting via absorptance Modulation☆21Oct 16, 2024Updated last year
- Deploy YOLOv8 in Unity using Sentis☆21Apr 20, 2024Updated 2 years ago
- ☆22Nov 3, 2025Updated 10 months ago
- A collection of various llm pruning implementations, training code for GPUs & TPUs, and evaluation script.☆75Apr 20, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A Novel Approach for Effective Multi-View Clustering with Information-Theoretic Perspective is a paper accepted by NeurIPS 2023☆11May 15, 2024Updated 2 years ago
- Cross-modal Clustering with Deep Correlated Information Bottleneck Method☆11Aug 7, 2022Updated 4 years ago
- Deriving steepest descent convergence bounds and hyperparameter scaling laws in machine learning optimization from first principles, form…☆17Apr 11, 2026Updated 5 months ago
- [ICLR 2025] Palu: Compressing KV-Cache with Low-Rank Projection☆163Feb 20, 2025Updated last year
- Pytorch Implementation of CLIP-Lite | Accepted at AISTATS 2023☆14Mar 17, 2023Updated 3 years ago
- KDD Cup 2022 Baidu Wind Power Forecast项目:百度风电功率预测赛 (Paddle Track 5th)☆13Jul 29, 2022Updated 4 years ago
- [AAAI 2023] Scalable Attributed-Graph Subspace Clustering☆14Jul 16, 2023Updated 3 years ago
- One command · One Microsoft login · Zero repeated auth Hours of uninterrupted access to NYU Torch from your terminal and IDE.☆16Jul 17, 2026Updated 2 months ago
- This is Official implementation for T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasonin…☆24Mar 5, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Least Squares Regression for subspace clustering☆11May 27, 2018Updated 8 years ago
- Official implementation of "ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppress…☆23Sep 25, 2025Updated 11 months ago
- Official Implementation (Pytorch) of the "Representation Shift: Unifying Token Compression with FlashAttention", ICCV 2025☆36Feb 22, 2026Updated 7 months ago
- Official Repo of CudaForge☆89Dec 2, 2025Updated 9 months ago
- [ICCV 2025] Official code of paper "Dynamic Multi-Layer Null Space Projection for Vision-Language Continual Learning"☆28Sep 8, 2025Updated last year
- [ACM MM 2026 Oral]⚡ZEUS accelerates your diffuser. Any modality. Any model. Any scheduler. https://yixiao-wang-stats.github.io/zeus/☆22Jun 2, 2026Updated 3 months ago
- Quantize transformers to any learned arbitrary 4-bit numeric format☆59Jul 2, 2026Updated 2 months ago