[NeurIPS 2023] Token-Scaled Logit Distillation for Ternary Weight Generative Language Models
☆18Dec 6, 2023Updated 2 years ago
Alternatives and similar repositories for TSLD
Users that are interested in TSLD are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TernGEMM: General Matrix Multiply Library with Ternary Weights for Fast DNN Inference☆14Feb 22, 2022Updated 4 years ago
- Layer-wise Pruning of Transformer Heads for Efficient Language Modeling☆22Feb 22, 2022Updated 4 years ago
- [NeurIPS 2023] SiT Dataset: Socially Interactive Pedestrian Trajectory Dataset for Social Navigation Robots☆81Oct 17, 2024Updated last year
- LLM Inference with Microscaling Format☆35Nov 12, 2024Updated last year
- AFPQ code implementation☆23Nov 6, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ACL 2024] A novel QAT with Self-Distillation framework to enhance ultra low-bit LLMs.☆140May 16, 2024Updated 2 years ago
- BESA is a differentiable weight pruning technique for large language models.☆17Mar 4, 2024Updated 2 years ago
- super-resolution; post-training quantization; model compression☆14Nov 10, 2023Updated 2 years ago
- ☆21May 4, 2026Updated 3 months ago
- ☆22Dec 5, 2022Updated 3 years ago
- PB-LLM: Partially Binarized Large Language Models☆158Nov 20, 2023Updated 2 years ago
- ☆11May 24, 2024Updated 2 years ago
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 10 months ago