☆35Dec 22, 2025Updated 9 months ago
Alternatives and similar repositories for quantized-training
Users that are interested in quantized-training are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequences☆33Mar 7, 2024Updated 2 years ago
- ☆26Apr 3, 2026Updated 5 months ago
- softfloat and softposit in Python☆15Aug 2, 2019Updated 7 years ago
- Tender: Accelerating Large Language Models via Tensor Decompostion and Runtime Requantization (ISCA'24)☆34Jul 4, 2024Updated 2 years ago
- [HPCA 2023] ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design☆134Jun 27, 2023Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- MICRO 2024 Evaluation Artifact for FuseMax☆18Aug 26, 2024Updated 2 years ago
- [ASPLOS 2026] M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization.☆17Jan 29, 2026Updated 7 months ago
- Here are some implementations of basic hardware units in RTL language (verilog for now), which can be used for area/power evaluation and …☆15Aug 25, 2023Updated 3 years ago
- Low Precision Arithmetic Simulation in PyTorch - extension for posit and beyond☆16Dec 9, 2025Updated 9 months ago
- ☆48Aug 23, 2021Updated 5 years ago
- ☆22Jul 30, 2024Updated 2 years ago
- LLM Inference with Microscaling Format☆35Nov 12, 2024Updated last year
- ☆38Jul 27, 2026Updated last month
- Training with Block Minifloat number representation☆18May 2, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆11Aug 2, 2024Updated 2 years ago
- [ACL 2026 Main] Code for the paper "ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs"☆32Updated this week
- [TCAD'23] AccelTran: A Sparsity-Aware Accelerator for Transformers☆62Nov 22, 2023Updated 2 years ago
- ☆16Nov 22, 2022Updated 3 years ago
- ViTALiTy (HPCA'23) Code Repository☆25Mar 13, 2023Updated 3 years ago
- An open-source parameterizable NPU generator with full-stack multi-target compilation stack for intelligent workloads.☆86Sep 29, 2025Updated 11 months ago
- [ICML 2024] When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models☆35Jun 12, 2024Updated 2 years ago
- ☆160Jul 19, 2025Updated last year
- DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators☆20Oct 10, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Serpens is an HBM FPGA accelerator for SpMV☆23Jul 26, 2024Updated 2 years ago
- A Deep Learning Framework for the Posit Number System☆31Aug 5, 2024Updated 2 years ago
- Torch2Chip (MLSys, 2024)☆56Apr 2, 2025Updated last year
- PositNN - Framework for training and inference with neural nets usings posits☆21Aug 17, 2026Updated last month
- A systolic array simulator for multi-cycle MACs and varying-byte words, with the paper accepted to HPCA 2022.☆85Nov 7, 2021Updated 4 years ago
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. 🚀 The official implementation of https://arx…☆32Feb 17, 2025Updated last year
- ☆13Jul 11, 2023Updated 3 years ago
- Official implementation for "Pruning Large Language Models with Semi-Structural Adaptive Sparse Training" (AAAI 2025)☆20Jul 1, 2025Updated last year
- ☆18Apr 7, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Official Pytorch Implementation of "Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity"☆81Jul 7, 2025Updated last year
- [DAC 2026] FlashFPS☆16Jun 1, 2026Updated 3 months ago
- [NeurIPS 2023] Token-Scaled Logit Distillation for Ternary Weight Generative Language Models☆18Dec 6, 2023Updated 2 years ago
- Reconfigurable Stream Network Architecture☆18Aug 13, 2026Updated last month
- ☆124Nov 17, 2023Updated 2 years ago
- Universal number Posit HDL Arithmetic Architecture generator☆72Jun 24, 2019Updated 7 years ago
- RTL implementation of Flex-DPE.☆119Feb 22, 2020Updated 6 years ago