GGUF implementation in C as a library and a tools CLI program
☆357May 16, 2026Updated 2 months ago
Alternatives and similar repositories for gguf-tools
Users that are interested in gguf-tools are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A small utility library for parsing GGUF file info☆30Jan 27, 2025Updated last year
- Some random tools for working with the GGUF file format☆32Nov 24, 2023Updated 2 years ago
- A fork of llama3.c used to do some R&D on inferencing☆23Dec 20, 2024Updated last year
- ggml implementation of BERT☆502Feb 23, 2024Updated 2 years ago
- GGUF parser for Go☆14Mar 8, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- CLIP inference in plain C/C++ with no extra dependencies☆565Jun 19, 2025Updated last year
- Inference of Mamba, Mamba2 and Mamba3 models in pure C☆203Mar 18, 2026Updated 4 months ago
- Fast neural codec compression and generation for audio waveforms☆231Dec 4, 2024Updated last year
- Julia interface to the 🤗 Hub☆19May 30, 2026Updated 2 months ago
- HC-256 Stream cipher in x86 assembly☆19Nov 14, 2017Updated 8 years ago
- Implementation of ModernBERT in MLX☆21Jan 7, 2026Updated 7 months ago
- A new city of code on a cosmopolitan foundation.☆21Mar 19, 2021Updated 5 years ago
- Suno AI's Bark model in C/C++ for fast text-to-speech generation☆866Nov 16, 2024Updated last year
- Tensor library for machine learning☆15,159Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Cross-platform binary launcher with Cosmopolitan libc☆34Apr 12, 2025Updated last year
- INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model☆1,581Mar 23, 2025Updated last year
- Node.js module providing inference APIs for large language models, with simple CLI.☆25Dec 7, 2024Updated last year
- Train text generation model with JavaScript.☆15Jul 14, 2024Updated 2 years ago
- First token cutoff sampling inference example☆30Jan 15, 2024Updated 2 years ago
- Code Llama GGUF Demo☆10Aug 28, 2023Updated 2 years ago
- Port of Microsoft's BioGPT in C/C++ using ggml☆87Feb 21, 2024Updated 2 years ago
- A collection of experiments related to LLM inference with llama.cpp/mlx☆40Updated this week
- ☆27Feb 26, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- run ollama & gguf easily with a single command☆53May 15, 2024Updated 2 years ago
- Flux 2 image generation model pure C inference☆1,980Feb 13, 2026Updated 6 months ago
- ☆137Nov 9, 2024Updated last year
- Visual bag of words for fast image matching☆25Apr 27, 2023Updated 3 years ago
- Inference Vision Transformer (ViT) in plain C/C++ with ggml☆319Apr 11, 2024Updated 2 years ago
- A minimalistic C++ Jinja templating engine for LLM chat templates☆223Sep 22, 2025Updated 10 months ago
- Temporary mail - Keep your real mailbox clean and secure. Temp Mail provides temporary, secure, anonymous, free, disposable email address…☆12Mar 17, 2023Updated 3 years ago
- Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++☆6,736Updated this week
- iterate quickly with llama.cpp hot reloading. use the llama.cpp bindings with bun.sh☆51Oct 30, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- lightweight, standalone C++ inference engine for Google's Gemma models.☆7,021Updated this week
- ☆68Aug 19, 2024Updated last year
- Thin wrapper around GGML to make life easier☆49Jul 26, 2026Updated 2 weeks ago
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- convert a saved pytorch model to gguf and generate as much corresponding ggml c code as possible☆15Dec 19, 2023Updated 2 years ago
- Load and run Llama from safetensors files in C☆16Oct 24, 2024Updated last year
- A simple MLX implementation for pretraining LLMs on Apple Silicon.☆85Aug 20, 2025Updated 11 months ago