LLaVA server (llama.cpp).
☆183Oct 20, 2023Updated 2 years ago
Alternatives and similar repositories for llava-cpp-server
Users that are interested in llava-cpp-server are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- iterate quickly with llama.cpp hot reloading. use the llama.cpp bindings with bun.sh☆51Oct 30, 2023Updated 2 years ago
- A simple "Be My Eyes" web app with a llama.cpp/llava backend☆495Nov 28, 2023Updated 2 years ago
- Port of Suno AI's Bark in C/C++ for fast inference☆55Apr 15, 2024Updated 2 years ago
- ☆1,275Oct 24, 2023Updated 2 years ago
- CLIP inference in plain C/C++ with no extra dependencies☆558Jun 19, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Suno AI's Bark model in C/C++ for fast text-to-speech generation☆861Nov 16, 2024Updated last year
- Port of Microsoft's BioGPT in C/C++ using ggml☆87Feb 21, 2024Updated 2 years ago
- Inference Vision Transformer (ViT) in plain C/C++ with ggml☆31Nov 23, 2023Updated 2 years ago
- Semantic emoji finder. Python/dash UI. Uses sentence transformer embeddings and duckdb☆20Sep 15, 2025Updated 8 months ago
- Inference of Large Multimodal Models in C/C++. LLaVA and others☆48Oct 1, 2023Updated 2 years ago
- Inference Vision Transformer (ViT) in plain C/C++ with ggml☆314Apr 11, 2024Updated 2 years ago
- The Codec 2 speech codec, compiled to WASM using Emscripten.☆13Apr 27, 2023Updated 3 years ago
- Fine-tuning, DPO, RLHF, RLAIF on LLMs - Qwen3, Zephyr 7B GPTQ with 4-Bit Quantization, Mistral-7B-GPTQ☆15Jul 5, 2025Updated 10 months ago
- Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++☆6,080Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- GPT-2 small trained on phi-like data☆68Feb 18, 2024Updated 2 years ago
- Demo python script app to interact with llama.cpp server using whisper API, microphone and webcam devices.☆47Nov 6, 2023Updated 2 years ago
- Friendly Terminal Assistant for Developers☆17Mar 23, 2024Updated 2 years ago
- Port of MiniGPT4 in C++ (4bit, 5bit, 6bit, 8bit, 16bit CPU inference with GGML)☆572Aug 8, 2023Updated 2 years ago
- Fine-tune mistral-7B on 3090s, a100s, h100s☆727Oct 11, 2023Updated 2 years ago
- A Javascript library (with Typescript types) to parse metadata of GGML based GGUF files.☆52Jul 30, 2024Updated last year
- The llama-cpp-agent framework is a tool designed for easy interaction with Large Language Models (LLMs). Allowing users to chat with LLM …☆632Mar 9, 2026Updated 2 months ago
- Tensor library for machine learning☆274Apr 23, 2023Updated 3 years ago
- Python bindings for llama.cpp☆10,312May 18, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆12Jan 25, 2023Updated 3 years ago
- transformer tokenizers (e.g. BERT tokenizer) in C++ (WIP)☆18Apr 7, 2022Updated 4 years ago
- [NeurIPS 2024] Self-Optimization Improves the Efficiency of Code Generation☆14May 10, 2025Updated last year
- LLM-based code completion engine☆192Jan 23, 2025Updated last year
- ☆62Jun 13, 2024Updated last year
- Apache Lucene/Solr Guide☆13Oct 14, 2021Updated 4 years ago
- ☆134Nov 24, 2023Updated 2 years ago
- Extracts structured data from unstructured input. Programming language agnostic. Uses llama.cpp☆45May 16, 2024Updated 2 years ago
- Web App to transcribe memos using Whisper AI.☆18Oct 23, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An implementation of Compositional Attention: Disentangling Search and Retrieval by MILA☆14Jun 1, 2022Updated 3 years ago
- Cheat at search with LLMs course materials☆41Updated this week
- Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.☆2,938Apr 14, 2026Updated last month
- This repo is for handling Question Answering, especially for Multi-hop Question Answering☆69Dec 20, 2023Updated 2 years ago
- GGML implementation of BERT model with Python bindings and quantization.☆57Feb 19, 2024Updated 2 years ago
- Python bindings for the Transformer models implemented in C/C++ using GGML library.☆1,887Jan 28, 2024Updated 2 years ago
- ☆15Sep 8, 2023Updated 2 years ago