Docs for GGUF quantization (unofficial)
☆501Jul 19, 2025Updated last year
Alternatives and similar repositories for gguf-docs
Users that are interested in gguf-docs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Evaluation framework for GGUF☆15Apr 2, 2026Updated 4 months ago
- Evaluating practical performance of local multi-turn conversational LLMs.☆19Aug 1, 2025Updated last year
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆25Sep 1, 2025Updated 11 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,021Updated this week
- ☆18Jul 1, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- a character-ai like UI for LLM☆10Dec 3, 2024Updated last year
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- A simple, easy-to-customize pipeline for local RAG evaluation. Starter prompts and metric definitions included.☆24Jan 14, 2026Updated 6 months ago
- Local Qwen3 LLM inference. One easy-to-understand file of C source with no dependencies.☆195Jul 5, 2025Updated last year
- A user-friendly GUI for llama.cpp — convert, quantize, and run GGUF models without touching the terminal.☆21Jun 9, 2026Updated 2 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆148Aug 1, 2026Updated last week
- ☆21Mar 22, 2026Updated 4 months ago
- ☆16Jul 23, 2026Updated 2 weeks ago
- Writing Tools, Apple's AI-inspired app, enchants Windows, enhancing your pen with AI LLMs. One hotkey press, system-wide, fixes grammar, …☆32Jun 23, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆1,829Updated this week
- ☆19Oct 18, 2025Updated 9 months ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- High-Performance Text Deduplication Toolkit☆61Aug 25, 2025Updated 11 months ago
- LLM Frontend in a single html file☆751Dec 27, 2025Updated 7 months ago
- Review/Check GGUF files and estimate the memory usage and maximum tokens per second.☆285Jul 13, 2026Updated 3 weeks ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,306Updated this week
- ☆42Feb 25, 2026Updated 5 months ago
- LLM inference in C/C++☆123,193Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,124Updated this week
- Mirror from gitlab☆11Jan 9, 2021Updated 5 years ago
- Thin wrapper around GGML to make life easier☆49Jul 26, 2026Updated 2 weeks ago
- This is a repository for collaboration on vision-based robot navigation.☆12Jul 28, 2022Updated 4 years ago
- Local-first agent runtime for MCP workflows with explicit trust controls, replayable runs, and built-in evals.☆32Jul 4, 2026Updated last month
- Transplants vocabulary between language models, enabling the creation of draft models for speculative decoding WITHOUT retraining.☆54Oct 29, 2025Updated 9 months ago
- ☆340May 15, 2026Updated 2 months ago
- Minimal web client for chatting and roleplay with AI characters☆27Aug 21, 2025Updated 11 months ago
- The local UI to run and train text and diffusion models, including Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.☆69,756Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆167Updated this week
- An OpenVoice-based voice cloning tool, single executable file (~14M), supporting multiple formats without dependencies on ffmpeg, Python,…☆49Jan 18, 2026Updated 6 months ago
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆47Jan 13, 2026Updated 6 months ago
- Tensor library for machine learning☆15,137Updated this week
- InferX: Inference as a Service Platform☆233Updated this week
- Enhancing LLMs with LoRA☆224Oct 20, 2025Updated 9 months ago
- The official API server for Exllama. OAI compatible, lightweight, and fast.☆1,298Updated this week