Accessible large language models via k-bit quantization for PyTorch.
β8,446Aug 27, 2026Updated this week
Alternatives and similar repositories for bitsandbytes
Users that are interested in bitsandbytes are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fast and memory-efficient exact attentionβ24,794Updated this week
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,604Updated this week
- QLoRA: Efficient Finetuning of Quantized LLMsβ10,999Jun 10, 2024Updated 2 years ago
- Hackable and optimized Transformers building blocks, supporting a composable construction.β10,544Aug 7, 2026Updated 3 weeks ago
- An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.β5,070Apr 11, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers".β2,362Mar 27, 2024Updated 2 years ago
- Transformer related optimization, including BERT, GPTβ6,447Mar 27, 2024Updated 2 years ago
- [MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Accelerationβ3,623Jul 17, 2025Updated last year
- Development repository for the Triton language and compilerβ20,031Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β43,018Updated this week
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,841Updated this week
- Large Language Model Text Generation Inferenceβ10,890Mar 21, 2026Updated 5 months ago
- Train transformer language models with reinforcement learning.β19,169Updated this week
- Ongoing research training transformer models at scaleβ17,652Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- AutoAWQ implements the AWQ algorithm for 4-bit quantization with a 2x speedup during inference. Documentation:β2,349May 11, 2025Updated last year
- TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizatβ¦β14,493Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMsβ90,314Updated this week
- Tensor library for machine learningβ15,249Updated this week
- A framework for few-shot evaluation of language models.β13,819Updated this week
- Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"β13,769Dec 17, 2024Updated last year
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,505Updated this week
- 4 bits quantization of LLaMA using GPTQβ3,073Jul 13, 2024Updated 2 years ago
- [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Modelsβ1,679Jul 12, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Instruct-tune LLaMA on consumer hardwareβ18,906Jul 29, 2024Updated 2 years ago
- SGLang is a high-performance serving framework for large language models and multimodal models.β32,623Updated this week
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,526May 1, 2026Updated 3 months ago
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,472Updated this week
- FlashInfer: Kernel Library for LLM Servingβ6,273Updated this week
- PyTorch native quantization for training and inferenceβ2,959Updated this week
- Tools for merging pretrained large language models.β7,321Jun 17, 2026Updated 2 months ago
- Running large language models on a single GPU for throughput-oriented scenarios.β9,352Oct 28, 2024Updated last year
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,196Updated this week
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- PyTorch extensions for high performance and large scale training.β3,407Apr 26, 2025Updated last year
- Go ahead and axolotl questionsβ12,415Updated this week
- [ICLR 2024] Efficient Streaming Language Models with Attention Sinksβ7,268Jul 11, 2024Updated 2 years ago
- Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.β6,248Aug 22, 2025Updated last year
- PyTorch native post-training libraryβ5,802Updated this week
- RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable)β¦β14,686Updated this week
- Code and documentation to train Stanford's Alpaca models, and generate the data.β30,246Jul 17, 2024Updated 2 years ago