TLLM_QMM strips the implementation of quantized kernels of Nvidia's TensorRT-LLM, removing NVInfer dependency and exposes ease of use Pytorch module. We modified the dequantation and weight preprocessing to align with popular quantization alogirthms such as AWQ and GPTQ, and combine them with new FP8 quantization.
☆16Jul 5, 2024Updated 2 years ago
Alternatives and similar repositories for TLLM_QMM
Users that are interested in TLLM_QMM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 中文语料:大量人工标注样本,非常有价值 !!!☆11Aug 15, 2019Updated 7 years ago
- Awesome Quantization Paper lists with Codes☆10Feb 24, 2021Updated 5 years ago
- No pain HTML parsing library.☆12Apr 2, 2018Updated 8 years ago
- ☆27Feb 17, 2025Updated last year
- ☆13Nov 28, 2014Updated 11 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Bagua tutorials.☆13Sep 4, 2022Updated 3 years ago
- Möbius Transformation for Fast Inner Product Search on Graph☆23Jun 3, 2021Updated 5 years ago
- ☆10Jun 28, 2019Updated 7 years ago
- PyTorch code for our paper "Progressive Binarization with Semi-Structured Pruning for LLMs"☆13Jul 11, 2026Updated last month
- ☆27Jan 8, 2024Updated 2 years ago
- 这是一款刷单平台的后台管理,主要针对商家,员工,订单,任务等进行一系列的管理☆11May 8, 2019Updated 7 years ago
- Convert the PyTorch MaskRCNN model using the coremltool☆10Feb 8, 2025Updated last year
- 一个非常高效的字符串匹配工具,支持正向/反向最大匹配分词和多模式字符串精确匹配☆16Jul 29, 2023Updated 3 years ago
- PyTorch implementation of "Deep Transferring Quantization" (ECCV2020)☆18Jun 22, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 放一些论文,简历之类的latex模板☆12Aug 4, 2026Updated 3 weeks ago
- ☆14Mar 21, 2020Updated 6 years ago
- 16 bit serial multiplier in SystemVerilog☆13Oct 13, 2018Updated 7 years ago
- Wallace and Dadda tree multiplier generator in vhdl and verilog☆14Mar 14, 2026Updated 5 months ago
- ☆12Mar 13, 2023Updated 3 years ago
- AFP is a hardware-friendly quantization framework for DNNs, which is contributed by Fangxin Liu and Wenbo Zhao.☆13Nov 8, 2021Updated 4 years ago
- Sample examples of how to call collective operation functions on multi-GPU environments. A simple example of using broadcast, reduce, all…☆36Aug 28, 2023Updated 3 years ago
- Quantize pytorch model, support post-training quantization and quantization aware training methods☆15Jun 15, 2023Updated 3 years ago
- Adaptive floating-point based numerical format for resilient deep learning☆14Apr 11, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ECE 5745 Tutorial 8: SRAM Generators☆19Mar 5, 2022Updated 4 years ago
- implement bert in pure c++☆37Apr 29, 2020Updated 6 years ago
- High performance NCCL plugin for Bagua.☆15Sep 15, 2021Updated 4 years ago
- ☆11Jun 6, 2023Updated 3 years ago
- A C++ port of karpathy/micrograd, a tiny scalar-valued autograd engine and a neural net library☆13Nov 24, 2023Updated 2 years ago
- the completion of CNNs by myself☆14Oct 8, 2015Updated 10 years ago
- A highly optimized LLM inference acceleration engine for Llama and its variants.☆908Mar 18, 2026Updated 5 months ago
- RocksDB made replicated using Robust Distributed System Nucleus (rDSN) (Delta Learning)☆16Sep 15, 2015Updated 10 years ago
- A C++-based RPC framework☆12Oct 28, 2021Updated 4 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- **curve_fit_utils** is a Python module containing useful tools for curve fitting☆18Dec 23, 2017Updated 8 years ago
- PyTorch implementation of "MLP-Mixer: An all-MLP Architecture for Vision" Tolstikhin et al. (2021)☆31May 13, 2021Updated 5 years ago
- ☆22Jul 11, 2023Updated 3 years ago
- 在苏剑林老师的代码上改了一下,改成了python3.6,基于膨胀卷积,字词混合向量,radam梯度优化算法,百度百科词向量的阅读理解模型☆24Aug 28, 2019Updated 7 years ago
- 2020语言与智能技术竞赛:关系抽取任务(https://aistudio.baidu.com/aistudio/competition/detail/31?lang=zh_CN)☆23May 19, 2020Updated 6 years ago
- 32 bit pipelined binary floating point adder using IEEE-754 Single Precision Format in Verilog☆18Aug 27, 2020Updated 6 years ago
- Optimized Parallel Tiled Approach to perform Matrix Multiplication by taking advantage of the lower latency, higher bandwidth shared memo…☆17Sep 24, 2017Updated 8 years ago