In this repository, we explore model compression for transformer architectures via quantization. We specifically explore quantization aware training of the linear layers and demonstrate the performance for 8 bits, 4 bits, 2 bits and 1 bit (binary) quantization.
☆24May 14, 2021Updated 5 years ago
Alternatives and similar repositories for Compressed-Transformers
Users that are interested in Compressed-Transformers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Oct 24, 2022Updated 3 years ago
- bitfusion verilog implementation☆13Feb 21, 2022Updated 4 years ago
- DeiT implementation for Q-ViT☆26Apr 21, 2025Updated last year
- 基于Point Transformers复现点云分割任务,并使用HAQ算法进行自动量化压缩,几乎不影响精度☆26Aug 25, 2022Updated 4 years ago
- Post-Training Quantization for Vision transformers.☆245Jul 19, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Meanflow and multilingual for F5-TTS model☆16Aug 23, 2025Updated last year
- SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration (Full Paper Accepted in FPGA'24)☆38Mar 12, 2026Updated 6 months ago
- HLS Custom-Precision Floating-Point Library☆13Nov 6, 2017Updated 8 years ago
- ☆13Jun 4, 2024Updated 2 years ago
- Generate an FPGA design for a TWN☆11Nov 4, 2019Updated 6 years ago
- HLS project modeling various sparse accelerators.☆12Jan 11, 2022Updated 4 years ago
- Comparing Audio Features for Unsupervised Sound Classification☆10Jun 22, 2022Updated 4 years ago
- PyTorch implementation of Towards Efficient Training for Neural Network Quantization☆16Jan 16, 2020Updated 6 years ago
- ☆11Aug 2, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 一个基于AXI接口的PL端卷积加速器,可由PS端调用☆12Apr 15, 2023Updated 3 years ago
- Pytorch implementation of the paper : A Global-local Attention Framework for Weakly Labelled Audio Tagging.☆13Feb 6, 2021Updated 5 years ago
- ☆13Jul 2, 2016Updated 10 years ago
- Classify modulation of signals☆16Jan 16, 2020Updated 6 years ago
- Pytorch implementation of the paper "Debiasing the Cloze Task in Sequential Recommendation with Bidirectional Transformers".☆12Jan 22, 2023Updated 3 years ago
- LLM4HWDesign Starting Toolkit☆20Oct 4, 2024Updated last year
- Verilog and matlab implementation of tanh using Cordic algorithm☆11Jun 5, 2020Updated 6 years ago
- ☆30Dec 5, 2023Updated 2 years ago
- This repository contains code that was used as an example of how to use Python to download part of the AudioSet dataset and use Tensorflo…☆13Aug 24, 2017Updated 9 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Pytorch implementation of the paper : Modeling Label Dependencies for Audio Tagging with Graph Convolutional Network☆15Sep 18, 2020Updated 6 years ago
- Quantize pytorch model, support post-training quantization and quantization aware training methods☆15Jun 15, 2023Updated 3 years ago
- A linear array of PEs with RISC-V ISA targeting extreme high frequency on Xilinx ZYNQ Ultrascale+, specificially for applications such as…☆14Jun 4, 2024Updated 2 years ago
- [AAAI 2023 Oral] Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training☆14Apr 19, 2023Updated 3 years ago
- ☆33May 17, 2024Updated 2 years ago
- ☆12Nov 24, 2023Updated 2 years ago
- NMT model with BERT in tensorflow 2.0☆20Jul 24, 2019Updated 7 years ago
- Workflow for Executing CNN Networks on Zynq Ultrascale+ with VITIS AI. Detailed analysis, configuration, and execution of Convolutional N…☆21Apr 8, 2024Updated 2 years ago
- ClusterGAN PyTorch implementation☆12Feb 24, 2020Updated 6 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- AFP is a hardware-friendly quantization framework for DNNs, which is contributed by Fangxin Liu and Wenbo Zhao.☆13Nov 8, 2021Updated 4 years ago
- ☆16Jan 21, 2024Updated 2 years ago
- ☆19Mar 17, 2021Updated 5 years ago
- ☆13Nov 25, 2022Updated 3 years ago
- Embedded hardware accelerator of multilayer perceptrons for lightweight machine learning☆18Dec 3, 2016Updated 9 years ago
- Implementation of a Quantized Transformer Model☆20Mar 20, 2019Updated 7 years ago
- A simplified version for DMC (Deep Multimodal Clustering for Unsupervised Audiovisual Learning)☆19May 27, 2020Updated 6 years ago