Accessible large language models via k-bit quantization for PyTorch.
β8,514Sep 7, 2026Updated last month
Alternatives and similar repositories for bitsandbytes
Users that are interested in bitsandbytes are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fast and memory-efficient exact attentionβ25,099Updated this week
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,766Updated this week
- QLoRA: Efficient Finetuning of Quantized LLMsβ11,034Jun 10, 2024Updated 2 years ago
- Hackable and optimized Transformers building blocks, supporting a composable construction.β10,557Sep 23, 2026Updated 2 weeks ago
- An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.β5,065Apr 11, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers".β2,381Mar 27, 2024Updated 2 years ago
- Transformer related optimization, including BERT, GPTβ6,457Mar 27, 2024Updated 2 years ago
- [MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Accelerationβ3,649Jul 17, 2025Updated last year
- Development repository for the Triton language and compilerβ20,319Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β43,206Updated this week
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,911Updated this week
- Large Language Model Text Generation Inferenceβ10,880Mar 21, 2026Updated 6 months ago
- Train transformer language models with reinforcement learning.β19,468Updated this week
- Ongoing research training transformer models at scaleβ18,080Updated this week
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- AutoAWQ implements the AWQ algorithm for 4-bit quantization with a 2x speedup during inference. Documentation:β2,346May 11, 2025Updated last year
- TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizatβ¦β14,775Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMsβ93,332Updated this week
- Tensor library for machine learningβ15,451Updated this week
- A framework for few-shot evaluation of language models.β14,145Sep 14, 2026Updated 3 weeks ago
- Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"β13,835Dec 17, 2024Updated last year
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,569Updated this week
- 4 bits quantization of LLaMA using GPTQβ3,070Jul 13, 2024Updated 2 years ago
- [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Modelsβ1,694Jul 12, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Instruct-tune LLaMA on consumer hardwareβ18,896Jul 29, 2024Updated 2 years ago
- SGLang is a high-performance serving framework for large language models and multimodal models.β36,835Updated this week
- An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.β39,551May 1, 2026Updated 5 months ago
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,498Updated this week
- FlashInfer: Kernel Library for LLM Servingβ6,556Updated this week
- PyTorch native quantization for training and inferenceβ2,993Updated this week
- Tools for merging pretrained large language models.β7,390Sep 12, 2026Updated 3 weeks ago
- Running large language models on a single GPU for throughput-oriented scenarios.β9,345Oct 28, 2024Updated last year
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,228Sep 21, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- PyTorch extensions for high performance and large scale training.β3,405Apr 26, 2025Updated last year
- Go ahead and axolotl questionsβ12,532Updated this week
- [ICLR 2024] Efficient Streaming Language Models with Attention Sinksβ7,263Jul 11, 2024Updated 2 years ago
- Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.β6,259Aug 22, 2025Updated last year
- PyTorch native post-training libraryβ5,811Sep 9, 2026Updated 3 weeks ago
- RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable)β¦β14,742Sep 28, 2026Updated last week
- Code and documentation to train Stanford's Alpaca models, and generate the data.β30,235Jul 17, 2024Updated 2 years ago