☆23May 12, 2026Updated 2 months ago
Alternatives and similar repositories for llamacpp-gfx-906-turbo
Users that are interested in llamacpp-gfx-906-turbo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆78Jun 23, 2026Updated last month
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆26Jul 2, 2026Updated last month
- llama.cpp-gfx906☆140Mar 22, 2026Updated 4 months ago
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆15Feb 21, 2026Updated 5 months ago
- Random AI notes for working with local models or playing around with random machine learning bits.☆61Jun 7, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 7 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆433Feb 20, 2026Updated 5 months ago
- Proxy for OpenAI☆16Sep 2, 2025Updated 11 months ago
- ☆19Aug 19, 2025Updated 11 months ago
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated 3 weeks ago
- A llamacpp wrapper to manage and monitor your llama server instance over a web ui.☆21Jun 16, 2026Updated last month
- A lightweight chat interface for interacting with local models, featuring persistent memory using a seamless SQLite database to store you…☆34Sep 15, 2025Updated 10 months ago
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆71May 4, 2025Updated last year
- A repository to store helpful information and emerging insights in regard to LLMs☆21Oct 27, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- RDNA-native LLM inference engine in Rust.☆497Updated this week
- Loader extension for tabbyAPI in SillyTavern☆27Jun 30, 2025Updated last year
- A premium RAG-based AI Assistant built with React and FastAPI. Features efficient document indexing and high-accuracy retrieval-augmented…☆19Jun 3, 2026Updated 2 months ago
- Structured local memory storage and retrieval for LLM agents☆15May 19, 2026Updated 2 months ago
- Neural-enhanced conversational memory for AI agents — 10 micro-networks, tri-hybrid storage, bio-inspired retrieval☆15Mar 27, 2026Updated 4 months ago
- Qwen 3.5 in C☆24Mar 28, 2026Updated 4 months ago
- A conversational voice-to-voice open-weights LLM-powered assistant designed to run on high-end consumer or workstation class hardware☆75Jul 6, 2026Updated last month
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆22Apr 3, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- the stanford graphbase, by donald knuth☆18Nov 26, 2015Updated 10 years ago
- Batch processor to enable large content be digested by Ollama, focused around book processing and translations by default, fully, configu…☆36Oct 27, 2025Updated 9 months ago
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 10 months ago
- Script Execution service☆13Nov 21, 2016Updated 9 years ago
- minimal fill-in-the-middle autocomplete for vscode/codium, for use with llama.cpp infill or any openai-compatible server☆16May 10, 2026Updated 2 months ago
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆167Updated this week
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆49Nov 14, 2025Updated 8 months ago
- Hindsight Cookbook - examples on how to use Hindsight☆36Jun 5, 2026Updated 2 months ago
- A desktop GUI for Flux 1.1 Pro built using DelphiFMX For Python☆11Oct 5, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆55Updated this week
- A UI for stable-diffusion.cpp.☆41Jul 22, 2026Updated 2 weeks ago
- dockerfiles, compose files, benchmarks for 2x R9700 serving☆33Updated this week
- Simple agent framework using Ollama tool calling☆10Aug 27, 2024Updated last year
- GFPGAN face reconstruction with ncnn on a bare Raspberry Pi☆14Jan 4, 2023Updated 3 years ago
- The Lily programming language ⚜☆11Updated this week
- LexiCrawler is a powerful Go-based web crawling API meticulously designed to extract, clean, and transform web page content into a pristi…☆48Feb 27, 2025Updated last year