Code for "RSQ: Learning from Important Tokens Leads to Better Quantized LLMs"
☆24Mar 25, 2026Updated 5 months ago
Alternatives and similar repositories for rsq
Users that are interested in rsq are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [COLM 2025] DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation; 知乎:https://zhuanlan.zhihu.c…☆30Mar 5, 2025Updated last year
- ☆35Mar 28, 2025Updated last year
- ☆25Oct 31, 2024Updated last year
- ☆12Oct 9, 2023Updated 2 years ago
- AFPQ code implementation☆23Nov 6, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR2025]: OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitt…☆95Apr 8, 2025Updated last year
- (EMNLP 2025 Main) RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives☆37Dec 20, 2025Updated 8 months ago
- ☆23Sep 19, 2024Updated last year
- ☆54Nov 5, 2024Updated last year
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆177Nov 26, 2025Updated 9 months ago
- This repository contains code for the MicroAdam paper.☆21Dec 14, 2024Updated last year
- [ICLR 2025] CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion☆56Jul 1, 2025Updated last year
- [ICML 2025] Official PyTorch implementation of "FlatQuant: Flatness Matters for LLM Quantization"☆227Nov 25, 2025Updated 9 months ago
- Code repo for "CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs".☆17Sep 15, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The official repository of NeurIPS'25 paper "Ada-R1: From Long-Cot to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization"☆24May 6, 2026Updated 3 months ago
- [ICML2024 Spotlight] Fine-Tuning Pre-trained Large Language Models Sparsely☆24Jun 26, 2024Updated 2 years ago
- Modified version of fairseq, including new implementations for criterions using reinforcement learning methods.☆11Aug 14, 2019Updated 7 years ago
- ☆18Feb 22, 2025Updated last year
- source code of the paper: Robust Quantization: One Model to Rule Them All☆42Mar 24, 2023Updated 3 years ago
- TVBench: Redesigning Video-Language Evaluation☆15Jun 9, 2025Updated last year
- ☆30May 22, 2026Updated 3 months ago
- Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)☆12Jun 20, 2025Updated last year
- Official implementation of "Can Test-Time Scaling Improve World Foundation Model?"☆15Jul 12, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- PyTorch code for full quantization of DNN using BCGD☆14Jul 24, 2019Updated 7 years ago
- ☆15Mar 21, 2025Updated last year
- [EMNLP 2025 Findings] MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation☆15Aug 22, 2025Updated last year
- ☆20Oct 13, 2024Updated last year
- Official code for DAM: Dynamic Adapter Merging for Continual Video QA Learning☆15Apr 25, 2024Updated 2 years ago
- ☆18Mar 4, 2024Updated 2 years ago
- ☆33Nov 11, 2024Updated last year
- Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"☆20Jan 18, 2026Updated 7 months ago
- [WWW 2026 Oral] MoE-CL:Self-Evolving LLMs via Continual Instruction Tuning☆22Dec 1, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆22Apr 27, 2026Updated 4 months ago
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆23May 25, 2026Updated 3 months ago
- A simple behavior that can be attached to a Page to display a custom TitleBar with a Full Screen Mode toggle. UWP only.☆12Aug 5, 2015Updated 11 years ago
- Artifact evaluation for HPCA'24 paper Lightening-Transformer: A Dynamically-operated Optically-interconnected Photonic Transformer Accele…☆11Mar 3, 2024Updated 2 years ago
- Unofficial Implementation of Consistency Models in Pytorch☆15Mar 18, 2023Updated 3 years ago
- [TCSVT] Regularity Learning via Explicit Distribution Modeling for Skeletal Video Anomaly Detection☆17Jul 22, 2023Updated 3 years ago
- Kinematics analytical solution and inverse solution for KUKA IIWA 7DOF robot.☆15Aug 22, 2026Updated last week