The llama-cpp-agent framework is a tool designed for easy interaction with Large Language Models (LLMs). Allowing users to chat with LLM models, execute structured function calls and get structured output. Works also with models not fine-tuned to JSON output and function calls.
☆652Mar 9, 2026Updated 5 months ago
Alternatives and similar repositories for llama-cpp-agent
Users that are interested in llama-cpp-agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ToolAgents is a lightweight and flexible framework for creating function-calling agents with various language models and APIs.☆38Apr 30, 2026Updated 3 months ago
- ☆32Dec 29, 2023Updated 2 years ago
- Locally running LLM with internet access☆96Jul 23, 2026Updated 2 weeks ago
- Python bindings for llama.cpp☆10,538Updated this week
- function calling-based LLM agents☆292Sep 16, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A Comprehensive survey on business use cases of AI that help them thrive in the digital economy☆13Oct 7, 2020Updated 5 years ago
- An experimental desktop client for using Claude Desktop's MCP with Novelcrafter codices.☆11Dec 3, 2024Updated last year
- Chat language model that can use tools and interpret the results☆1,595Jun 30, 2026Updated last month
- TypeScript generator for llama.cpp Grammar directly from TypeScript interfaces☆146Jul 9, 2024Updated 2 years ago
- This GUI aims to simplify the process of converting GGUF files to llamafile format by providing an intuitive and convenient way for users…☆14Jan 2, 2026Updated 7 months ago
- Inference of Mamba, Mamba2 and Mamba3 models in pure C☆203Mar 18, 2026Updated 4 months ago
- A guidance compatibility layer for llama-cpp-python☆37Sep 11, 2023Updated 2 years ago
- Harness LLMs with Multi-Agent Programming☆4,090Jul 29, 2026Updated last week
- Your Trusty Memory-enabled AI Companion - Simple RAG chatbot optimized for local LLMs | 12 Languages Supported | OpenAI API Compatible☆352Feb 28, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Python bindings for the Transformer models implemented in C/C++ using GGML library.☆1,885Jan 28, 2024Updated 2 years ago
- Create Custom LLMs☆1,859Jun 27, 2026Updated last month
- Enforce the output format (JSON Schema, Regex etc) of a language model☆2,027Apr 4, 2026Updated 4 months ago
- A multimodal, function calling powered LLM webui.☆213Sep 23, 2024Updated last year
- Simple agent framework using Ollama tool calling☆10Aug 27, 2024Updated last year
- cli tool to quantize gguf, gptq, awq, hqq and exl2 models☆78Dec 17, 2024Updated last year
- ☆352Mar 5, 2026Updated 5 months ago
- Converts JSON-Schema to GBNF grammar to use with llama.cpp☆56Nov 27, 2023Updated 2 years ago
- ☆137Jun 30, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆19Jun 5, 2023Updated 3 years ago
- A tool for generating function arguments and choosing what function to call with local LLMs☆435Mar 12, 2024Updated 2 years ago
- An AI assistant beyond the chat box.☆329Mar 11, 2024Updated 2 years ago
- ☆38Mar 12, 2024Updated 2 years ago
- WilmerAI is one of the oldest LLM semantic routers. It uses multi-layer prompt routing and complex workflows to allow you to not only cre…☆826Updated this week
- The official API server for Exllama. OAI compatible, lightweight, and fast.☆1,298Updated this week
- A library for working with GBNF files☆31May 27, 2026Updated 2 months ago
- A frontend for creative writing with LLMs☆170Jul 15, 2024Updated 2 years ago
- A simple experiment on letting two local LLM have a conversation about anything!☆112Jul 3, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Tools for merging pretrained large language models.☆7,288Jun 17, 2026Updated last month
- ☆1,437Dec 22, 2025Updated 7 months ago
- An easy-to-understand framework for LLM samplers that rewind and revise generated tokens☆151Jan 7, 2026Updated 7 months ago
- A fast inference library for running LLMs locally on modern consumer-class GPUs☆4,602Mar 4, 2026Updated 5 months ago
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- Efficient visual programming for AI language models☆356May 13, 2025Updated last year
- Like grep but for natural language questions. Based on Mistral 7B or Mixtral 8x7B.☆386Mar 13, 2024Updated 2 years ago