Due to the huge vocaburary size (151,936) of Qwen models, the Embedding and LM Head weights are excessively heavy. Therefore, this project provides a Tokenizer vocabulary shearing solution for Qwen and Qwen-VL.
☆41Jan 6, 2026Updated 7 months ago
Alternatives and similar repositories for Qwen-Tokenizer-Pruner
Users that are interested in Qwen-Tokenizer-Pruner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codebase for Instruction Following without Instruction Tuning☆36Sep 24, 2024Updated last year
- Software kit for Qualcomm Cloud AI 100☆19Dec 15, 2025Updated 8 months ago
- ☆309Apr 6, 2023Updated 3 years ago
- Structured Neuron Level Pruning to compress Transformer-based models [ECCV'24]☆16Aug 7, 2024Updated 2 years ago
- Các thí nghiệm liên quan tới LLMs cho tiếng Việt (insprised by Physics of LLMs Series)☆11Oct 21, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- GPTQ inference TVM kernel☆41Apr 25, 2024Updated 2 years ago
- super-resolution; post-training quantization; model compression☆14Nov 10, 2023Updated 2 years ago
- Interactive Temporal Consistency for Video Streams (Computer Graphics Forum, 2023)☆13Jan 27, 2025Updated last year
- ☆17Apr 10, 2025Updated last year
- A Pytorch implementation of Pensieve (SIGCOMM'18)☆12Jun 17, 2020Updated 6 years ago
- Conic10K: A large-scale dataset for closed-vocabulary math problem understanding. Accepted to EMNLP2023 Findings.☆33Dec 6, 2023Updated 2 years ago
- ☆15Oct 3, 2024Updated last year
- ☆14Oct 21, 2024Updated last year
- This is the repository for paper EscapeBench: Pushing Language Models to Think Outside the Box☆18Dec 19, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts☆36Jul 2, 2024Updated 2 years ago
- Streamable Text-to-Speech model using a language modeling approach, without vector quantization☆108May 20, 2025Updated last year
- ☆30Dec 27, 2024Updated last year
- [NeurIPS 2023] LLM-Pruner: On the Structural Pruning of Large Language Models. Support Llama-3/3.1, Llama-2, LLaMA, BLOOM, Vicuna, Baich…☆1,134Oct 7, 2024Updated last year
- Codes for ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding [ICML 2025]]☆51Jul 22, 2025Updated last year
- [ICLR 2024] Jaiswal, A., Gan, Z., Du, X., Zhang, B., Wang, Z., & Yang, Y. Compressing llms: The truth is rarely pure and never simple.☆27Apr 21, 2025Updated last year
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆37Aug 14, 2024Updated 2 years ago
- Code and model for paper <Mutual Information Maximization for Effective Lip Reading>☆19Sep 4, 2020Updated 5 years ago
- This is the official repository for Vista dataset - A Vietnamese multimodal dataset contains more than 700,000 samples of conversations a…☆26May 14, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Top Picks for Data Science Self-Study: From Newbies to Pros!☆11Apr 2, 2024Updated 2 years ago
- ☆12Dec 30, 2020Updated 5 years ago
- This is the official implementation to the EMNLP 2024 paper: Modeling Layout Reading Order as Ordering Relations for Visually-rich Docume…☆32Jan 19, 2026Updated 6 months ago
- ☆65Jun 12, 2025Updated last year
- [CVPR 2023] PD-Quant: Post-Training Quantization Based on Prediction Difference Metric☆61Mar 23, 2023Updated 3 years ago
- LaTeX Beamer template crafted for University of Illinois Chicago☆12Dec 7, 2024Updated last year
- a simple programming language under development☆11Dec 3, 2023Updated 2 years ago
- code for ml2023-spring-hung-yi-lee(李宏毅)☆13Apr 10, 2023Updated 3 years ago
- [NeurIPS 2023] Generalized Logit Adjustment☆40Apr 21, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆13Nov 5, 2024Updated last year
- a pytorch implementation of pensieve (https://github.com/hongzimao/pensieve)☆21Dec 24, 2019Updated 6 years ago
- RACE is a multi-dimensional benchmark for code generation that focuses on Readability, mAintainability, Correctness, and Efficiency.☆14Oct 12, 2024Updated last year
- ☆63May 19, 2025Updated last year
- ☆15Feb 11, 2025Updated last year
- [ICLR 2025] Causal Graphical Models for Vision-Language Compositional Understanding☆10Apr 15, 2025Updated last year
- ☆45Nov 1, 2025Updated 9 months ago