Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
☆291Aug 31, 2026Updated this week
Alternatives and similar repositories for gguf-parser-go
Users that are interested in gguf-parser-go are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LM inference server implementation based on *.cpp.☆292Nov 24, 2025Updated 9 months ago
- Download models from the Ollama library, without Ollama☆151Aug 12, 2026Updated 2 weeks ago
- A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.☆5,575Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,160Updated this week
- Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)☆921Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Rancher Prime GC Catalog☆13Updated this week
- Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d,…☆1,665Jul 19, 2026Updated last month
- LLM inference in C/C++☆23Oct 4, 2024Updated last year
- automatically quant GGUF models☆227Dec 23, 2025Updated 8 months ago
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- A simple, yet highly customisable DPadView for Android.☆12Mar 18, 2020Updated 6 years ago
- VellumForge2 is a Golang CLI for generating high-quality Direct Preference Optimization datasets via a hierarchical prompt pipeline with …☆21Feb 15, 2026Updated 6 months ago
- ☆24Dec 29, 2025Updated 8 months ago
- Simple model memory requirements calculator for GGUF☆85Jan 20, 2026Updated 7 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Editor with LLM generation tree exploration☆87Feb 12, 2025Updated last year
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,527Updated this week
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆155Aug 21, 2026Updated last week
- ☆21Jun 8, 2025Updated last year
- Docker/podman container for llama.cpp/vllm/exllamav{2,3} orchestrated using llama-swap☆18Jun 17, 2026Updated 2 months ago
- LLama.cpp golang bindings☆937Updated this week
- Create text chunks which end at natural stopping points without using a tokenizer☆26Nov 26, 2025Updated 9 months ago
- A refeference of text models that can be used in the AI Horde☆13Updated this week
- WebAssembly binding for llama.cpp - Enabling on-browser LLM inference☆1,189Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Accepts a Hugging Face model URL, automatically downloads and quantizes it using Bits and Bytes.☆38Mar 12, 2024Updated 2 years ago
- ☆18Jul 1, 2025Updated last year
- Tensor library for machine learning☆15,267Updated this week
- ☆137Nov 9, 2024Updated last year
- A go wrapper around the rwkv.cpp library☆20Mar 4, 2024Updated 2 years ago
- d.run website☆18Aug 25, 2026Updated last week
- VS Code extension for LLM-assisted code/text completion☆1,495Aug 19, 2026Updated last week
- SPLAA is an AI assistant framework that utilizes voice recognition, text-to-speech, and tool-calling capabilities to provide a conversati…☆29May 6, 2025Updated last year
- Automated .NET SDKs for your APIs☆90Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Automated project for generating passive income by sharing your internet connection. Compatible with platforms like Honeygain, PacketStre…☆16Jul 25, 2026Updated last month
- AI Assistant☆21Feb 21, 2026Updated 6 months ago
- LLM inference in C/C++☆20Updated this week
- An unofficial collection of precompiled WebP binaries for all of Apple's current platforms.☆20Apr 19, 2022Updated 4 years ago
- llama.cpp gguf file parser for javascript☆51Dec 11, 2024Updated last year
- Forked from ggerganov/llama.cpp☆18Updated this week
- Tiny Llama model trained to play chess☆31Jul 22, 2025Updated last year