Neural network model repository for highly sparse and sparse-quantized models with matching sparsification recipes
☆388Jun 2, 2025Updated last year
Alternatives and similar repositories for sparsezoo
Users that are interested in sparsezoo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Top-level directory for documentation and general content☆120Jun 2, 2025Updated last year
- ML model optimization product to accelerate inference.☆325Jun 2, 2025Updated last year
- Libraries for applying sparsification recipes to neural networks with a few lines of code, enabling faster and smaller models☆2,144Jun 2, 2025Updated last year
- Sparsity-aware deep learning inference runtime for CPUs☆3,154Jun 2, 2025Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMs☆267Dec 4, 2025Updated 10 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A simple GPU reservation tool for single host shared development systems☆31Aug 24, 2026Updated last month
- Pytorch distributed backend extension with compression support☆17Mar 24, 2025Updated last year
- McPAT modeling framework☆13Oct 18, 2014Updated 11 years ago
- Awesome Quantization Paper lists with Codes☆10Feb 24, 2021Updated 5 years ago
- MAML implementation with pytorch☆11Sep 23, 2020Updated 6 years ago
- A model compression and acceleration toolbox based on pytorch.☆333Jan 12, 2024Updated 2 years ago
- Refine high-quality datasets and visual AI models☆11,165Updated this week
- [TCAD 2021] Block Convolution: Towards Memory-Efficient Inference of Large-Scale CNNs on FPGA☆17Jul 7, 2022Updated 4 years ago
- The official PyTorch implementation of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024 paper Hyp²Nav:…☆18Oct 29, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Run zero-shot prediction models on your data☆37Dec 19, 2024Updated last year
- PyTorch implementation for the APoT quantization (ICLR 2020)☆287Dec 11, 2024Updated last year
- Video-Text Representation Learning via Differentiable Weak Temporal Alignment (CVPR 2022)☆19Apr 19, 2024Updated 2 years ago
- ☆11Jan 4, 2022Updated 4 years ago
- We have implemented a framework that supports developers to structured prune neural networks of Tensorflow Models☆29Nov 7, 2024Updated last year
- Code accompanying the NeurIPS 2020 paper: WoodFisher (Singh & Alistarh, 2020)☆54Mar 8, 2021Updated 5 years ago
- ☆10Jul 27, 2020Updated 6 years ago
- Benchmark PyTorch Custom Operators☆14Jul 6, 2023Updated 3 years ago
- Simulation and Optimization Library for Integrated Optics in Julia.☆11Sep 27, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Quantize pytorch model, support post-training quantization and quantization aware training methods☆15Jun 15, 2023Updated 3 years ago
- Magicube is a high-performance library for quantized sparse matrix operations (SpMM and SDDMM) of deep learning on Tensor Cores.☆92Nov 23, 2022Updated 3 years ago
- Code for the ICML 2023 paper "SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot".☆900Aug 20, 2024Updated 2 years ago
- ☆24Apr 20, 2024Updated 2 years ago
- Code for the paper "QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models".☆278Nov 3, 2023Updated 2 years ago
- ☆18Feb 7, 2024Updated 2 years ago
- yolov5 pruning (SFP Pruning、Nework Slimming)☆19Oct 5, 2021Updated 5 years ago
- MobileSAM のエンコーダー/デコーダーをONNXに変換し、推論するサンプル☆12Apr 11, 2024Updated 2 years ago
- structured sparsity regularization☆14Oct 12, 2019Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tools for simple inference testing using TensorRT, CUDA and OpenVINO CPU/GPU and CPU providers. Simple Inference Test for ONNX.☆25Sep 7, 2025Updated last year
- MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models (CVPR 2023)☆35Apr 23, 2024Updated 2 years ago
- Adding new tasks to T0 without catastrophic forgetting☆33Oct 20, 2022Updated 3 years ago
- [NeurIPS 2023] ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer☆31Dec 6, 2023Updated 2 years ago
- RTL implementation of Flex-DPE.☆119Feb 22, 2020Updated 6 years ago
- Extracting six domain-specific QA datasets from MS MARCO☆17Dec 1, 2019Updated 6 years ago
- Official implementation of "Searching for Winograd-aware Quantized Networks" (MLSys'20)☆27Oct 3, 2023Updated 3 years ago