Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
☆285Jul 13, 2026Updated 3 weeks ago
Alternatives and similar repositories for gguf-parser-go
Users that are interested in gguf-parser-go are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LM inference server implementation based on *.cpp.☆292Nov 24, 2025Updated 8 months ago
- Download models from the Ollama library, without Ollama☆150Nov 13, 2024Updated last year
- A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.☆5,458Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,021Updated this week
- Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d,…☆1,649Jul 19, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A text-to-speech and speech-to-text server compatible with the OpenAI API, supporting Whisper, FunASR, Bark, and CosyVoice backends.☆215Jul 17, 2026Updated 3 weeks ago
- Provides a unified interface to detect GPU resources and manages GPU workloads.☆16Updated this week
- LLM inference in C/C++☆23Oct 4, 2024Updated last year
- automatically quant GGUF models☆226Dec 23, 2025Updated 7 months ago
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- 🤖 AI-powered CLI for file reorganization. Runs fully locally — no data leaves your machine.☆20Jul 2, 2025Updated last year
- VellumForge2 is a Golang CLI for generating high-quality Direct Preference Optimization datasets via a hierarchical prompt pipeline with …☆21Feb 15, 2026Updated 5 months ago
- ☆24Dec 29, 2025Updated 7 months ago
- Deliver LLMs of GGUF format via Dockerfile.☆15Oct 24, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Simple model memory requirements calculator for GGUF☆84Jan 20, 2026Updated 6 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,306Updated this week
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆148Aug 1, 2026Updated last week
- ☆21Jun 8, 2025Updated last year
- Docker/podman container for llama.cpp/vllm/exllamav{2,3} orchestrated using llama-swap☆18Jun 17, 2026Updated last month
- 6800% faster "os" module replacement. A drop-in replacement for Python's standard 'OS' module. Fully-rewritten, optimized, and speeded-up…☆18Feb 20, 2023Updated 3 years ago
- LLama.cpp golang bindings☆931Jul 20, 2026Updated 3 weeks ago
- Create text chunks which end at natural stopping points without using a tokenizer☆26Nov 26, 2025Updated 8 months ago
- A refeference of text models that can be used in the AI Horde☆13Jul 18, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- WebAssembly binding for llama.cpp - Enabling on-browser LLM inference☆1,161Jun 17, 2026Updated last month
- Accepts a Hugging Face model URL, automatically downloads and quantizes it using Bits and Bytes.☆38Mar 12, 2024Updated 2 years ago
- ☆18Jul 1, 2025Updated last year
- Tensor library for machine learning☆15,137Updated this week
- ☆137Nov 9, 2024Updated last year
- A go wrapper around the rwkv.cpp library☆20Mar 4, 2024Updated 2 years ago
- d.run website☆18Updated this week
- VS Code extension for LLM-assisted code/text completion☆1,477Updated this week
- SPLAA is an AI assistant framework that utilizes voice recognition, text-to-speech, and tool-calling capabilities to provide a conversati…☆29May 6, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Collection of Dockerfiles to build images for various inference services across different accelerated backends.☆15Jul 15, 2026Updated 3 weeks ago
- Automated project for generating passive income by sharing your internet connection. Compatible with platforms like Honeygain, PacketStre…☆15Jul 25, 2026Updated 2 weeks ago
- AI Assistant☆21Feb 21, 2026Updated 5 months ago
- Benchmark structured generation libraries☆31Oct 25, 2024Updated last year
- LLM inference in C/C++☆18Updated this week
- llama.cpp gguf file parser for javascript☆50Dec 11, 2024Updated last year
- Forked from ggerganov/llama.cpp☆17Updated this week