This repository provides the official implementation of QSVD, a method for efficient low-rank approximation that unifies Query-Key-Value (QKV) weight compression in low-precision Vision-Language Models (VLMs).
☆28May 16, 2026Updated 2 months ago
Alternatives and similar repositories for QSVD
Users that are interested in QSVD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] Dobi-SVD : Differentiable SVD for LLM Compression and Some New Perspectives"☆54Oct 19, 2025Updated 9 months ago
- [AAAI 2026] Official implementation of "FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models". If you find this reposi…☆17May 1, 2026Updated 3 months ago
- [NeurIPS 2025 (spotlight)] HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs☆16Dec 17, 2025Updated 7 months ago
- [ICLR 2025🔥] SVD-LLM & [NAACL 2025🔥] SVD-LLM V2☆303Aug 28, 2025Updated 11 months ago
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization☆44Aug 13, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆51May 9, 2026Updated 3 months ago
- [ICLR 2026 Oral] Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation☆104May 8, 2026Updated 3 months ago
- TA's implementation for the project of Computer Architecture and Intelligent Chip Design (23 Spring)☆10May 20, 2023Updated 3 years ago
- Research in compressing convolutional layers of CNN using low-rank Tucker tensor decomposition☆12Nov 1, 2023Updated 2 years ago
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated last year
- ☆30Feb 6, 2026Updated 6 months ago
- PyTorch code for our paper "AdaSVD: Adaptive Singular Value Decomposition for Large Language Models"☆15Mar 9, 2025Updated last year
- 6 DoF Directional Room Impulse Response (RIR) with Dense Loudspeaker Grid☆17Aug 31, 2023Updated 2 years ago
- Code implementation of GPTAQ (https://arxiv.org/abs/2504.02692)☆94Jul 28, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆37Dec 24, 2025Updated 7 months ago
- UniQL official repository (ICLR 2026)☆17Jan 27, 2026Updated 6 months ago
- [ICRA 2025] Fast Global Localization on Neural Radiance Field☆17Jun 30, 2025Updated last year
- ☆16Oct 28, 2025Updated 9 months ago
- [ICML 2025] Official PyTorch implementation of "FlatQuant: Flatness Matters for LLM Quantization"☆225Nov 25, 2025Updated 8 months ago
- DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails☆34Feb 26, 2025Updated last year
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 3 months ago
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆18Nov 1, 2025Updated 9 months ago
- LLGS: Illuminating Gaussian Splatting via absorptance Modulation☆21Oct 16, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A collection of various llm pruning implementations, training code for GPUs & TPUs, and evaluation script.☆70Apr 20, 2026Updated 3 months ago
- ☆22Nov 3, 2025Updated 9 months ago
- A Novel Approach for Effective Multi-View Clustering with Information-Theoretic Perspective is a paper accepted by NeurIPS 2023☆11May 15, 2024Updated 2 years ago
- CastleHill: Separable Causal Diffusion / Varitaion Flow Maps for LTX-2 long-form video generation☆15May 19, 2026Updated 2 months ago
- Cross-modal Clustering with Deep Correlated Information Bottleneck Method☆11Aug 7, 2022Updated 4 years ago
- [ICLR 2025] Palu: Compressing KV-Cache with Low-Rank Projection☆161Feb 20, 2025Updated last year
- KDD Cup 2022 Baidu Wind Power Forecast项目:百度风电功率预测赛 (Paddle Track 5th)☆13Jul 29, 2022Updated 4 years ago
- One command · One Microsoft login · Zero repeated auth Hours of uninterrupted access to NYU Torch from your terminal and IDE.☆15Jul 17, 2026Updated 3 weeks ago
- This is Official implementation for T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasonin…☆24Mar 5, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ACM MM 2026]⚡ZEUS accelerates your diffuser. Any modality. Any model. Any scheduler. https://yixiao-wang-stats.github.io/zeus/☆20Jun 2, 2026Updated 2 months ago
- ☆13Jun 12, 2025Updated last year
- matlab code for hyperspectral target/anomaly detection