TLLM_QMM strips the implementation of quantized kernels of Nvidia's TensorRT-LLM, removing NVInfer dependency and exposes ease of use Pytorch module. We modified the dequantation and weight preprocessing to align with popular quantization alogirthms such as AWQ and GPTQ, and combine them with new FP8 quantization.
☆16Jul 5, 2024Updated 2 years ago
Alternatives and similar repositories for TLLM_QMM
Users that are interested in TLLM_QMM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆28Jul 6, 2026Updated last month
- ☆13Jun 4, 2024Updated 2 years ago
- Adapted iPerf3 iOS sample☆12Mar 15, 2017Updated 9 years ago
- Möbius Transformation for Fast Inner Product Search on Graph☆23Jun 3, 2021Updated 5 years ago
- ☆11Mar 18, 2019Updated 7 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆10Jun 28, 2019Updated 7 years ago
- PyTorch code for our paper "Progressive Binarization with Semi-Structured Pruning for LLMs"☆13Jul 11, 2026Updated last month
- ☆27Jan 8, 2024Updated 2 years ago
- An operator for managing Alluxio system on Kubernetes cluster☆13Jan 9, 2024Updated 2 years ago
- Deep Variational Information Bottleneck (DVIB) in PyTorch.☆10Apr 25, 2020Updated 6 years ago
- ☆10Mar 6, 2016Updated 10 years ago
- Convert the PyTorch MaskRCNN model using the coremltool☆10Feb 8, 2025Updated last year
- TensorRT-in-Action 是一个 GitHub 代码库,提供了使用 TensorRT 的代码示例,并有对应 Jupyter Notebook。☆15Jun 1, 2023Updated 3 years ago
- PyTorch implementation of "Deep Transferring Quantization" (ECCV2020)☆18Jun 22, 2022Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Low Precision Arithmetic Simulation in PyTorch - extension for posit and beyond☆16Dec 9, 2025Updated 8 months ago
- 放一些论文,简历之类的latex模板☆12Aug 4, 2026Updated last week
- Wallace and Dadda tree multiplier generator in vhdl and verilog☆14Mar 14, 2026Updated 5 months ago
- ☆12Mar 13, 2023Updated 3 years ago
- AFP is a hardware-friendly quantization framework for DNNs, which is contributed by Fangxin Liu and Wenbo Zhao.☆13Nov 8, 2021Updated 4 years ago
- Rate limiting library for python☆16Aug 6, 2023Updated 3 years ago
- Quantize pytorch model, support post-training quantization and quantization aware training methods☆15Jun 15, 2023Updated 3 years ago
- Adaptive floating-point based numerical format for resilient deep learning☆14Apr 11, 2022Updated 4 years ago
- MySQL GUI client for Ubuntu inspired by Sequel Pro☆21Aug 2, 2020Updated 6 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- PyTorch to CoreML: Writing custom layers with Metal shaders - torch.nn.functional.grid_sample operation☆16Jun 4, 2024Updated 2 years ago
- implement bert in pure c++☆37Apr 29, 2020Updated 6 years ago
- ☆11Jun 6, 2023Updated 3 years ago
- C library for handling proxy autoconfiguration (PAC) files.☆18Mar 22, 2018Updated 8 years ago
- A c++ hash map/table which utilizes simd (specifically Intel x86 SSE/AVX)☆12Apr 30, 2019Updated 7 years ago
- A highly optimized LLM inference acceleration engine for Llama and its variants.☆908Mar 18, 2026Updated 4 months ago
- Holistic 3D Human and Scene Mesh Estimation from Single View Images☆23Jun 14, 2021Updated 5 years ago
- A C++-based RPC framework☆12Oct 28, 2021Updated 4 years ago
- **curve_fit_utils** is a Python module containing useful tools for curve fitting☆18Dec 23, 2017Updated 8 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Jan 31, 2016Updated 10 years ago
- 32 bit pipelined binary floating point adder using IEEE-754 Single Precision Format in Verilog☆18Aug 27, 2020Updated 5 years ago
- ☆17Jan 3, 2025Updated last year
- OneFlow Serving☆20Apr 10, 2025Updated last year
- This project is for dealing with dynamic multiobjective optimization problems using a Multiobjective Evolutionary Algorithm.☆21May 6, 2018Updated 8 years ago
- Foolbox implementation for NeurIPS 2021 Paper: "Fast Minimum-norm Adversarial Attacks through Adaptive Norm Constraints".☆25Mar 16, 2022Updated 4 years ago
- ☆22Jul 5, 2026Updated last month