Pre-built wheels for llama-cpp-python across platforms and CUDA versions
☆84Apr 18, 2026Updated 5 months ago
Alternatives and similar repositories for llama-cpp-python-wheels
Users that are interested in llama-cpp-python-wheels are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-t…☆80Jan 19, 2026Updated 8 months ago
- Efficient Python bindings for the llama.cpp library☆552Updated this week
- 本地调用各种llama模型的comfyui节点,包含gemma4和Qwen3.5等常用模型☆53Aug 23, 2026Updated last month
- Quantized Attention that achieves speedups of 2.1-3.1x and 2.7-5.1x compared to FlashAttention2 and xformers, respectively, without lossi…☆157Updated this week
- Run gguf LLM models in Latest Version TextGen-webui and koboldcpp☆20Aug 6, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Fork of SpargeAttention (SparseSageAttention) for Windows wheels and easy installation☆40May 14, 2026Updated 4 months ago
- Pre-compiled Python whl for Flash-attention, SageAttention, NATTEN, xFormer etc