☆23May 12, 2026Updated 3 months ago
Alternatives and similar repositories for llamacpp-gfx-906-turbo
Users that are interested in llamacpp-gfx-906-turbo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆83Jun 23, 2026Updated 2 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆26Updated this week
- llama.cpp-gfx906☆139Updated this week
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆16Feb 21, 2026Updated 6 months ago
- Random AI notes for working with local models or playing around with random machine learning bits.☆63Jun 7, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 8 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆434Feb 20, 2026Updated 6 months ago
- ☆19Aug 19, 2025Updated last year
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated last month
- A llamacpp wrapper to manage and monitor your llama server instance over a web ui.☆22Jun 16, 2026Updated 2 months ago
- A lightweight chat interface for interacting with local models, featuring persistent memory using a seamless SQLite database to store you…☆34Sep 15, 2025Updated 11 months ago
- A LibGDX based Parallel AI Chess Game playable on many devices from Level 1 to Level 10☆14Apr 27, 2022Updated 4 years ago
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆71May 4, 2025Updated last year
- Easy access to unsaved files for vscode.☆12Jul 10, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The Offical AvdanOS demo☆16Aug 8, 2023Updated 3 years ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- (beta) Hindsight integration for pi coding agent (with queueing, past session ingestion, and a focus on best practices)☆23Aug 11, 2026Updated 2 weeks ago
- RDNA-native LLM inference engine in Rust.☆554Updated this week
- A simple Pong game in Flutter☆10Oct 19, 2019Updated 6 years ago
- Gives agents a real browser. URL in, pruned snapshot out. Replaces Playwright, Selenium, Puppeteer. Zero deps, zero wasted tokens.☆45Jul 13, 2026Updated last month
- ☆15Feb 10, 2020Updated 6 years ago
- Loader extension for tabbyAPI in SillyTavern☆27Jun 30, 2025Updated last year
- A premium RAG-based AI Assistant built with React and FastAPI. Features efficient document indexing and high-accuracy retrieval-augmented…☆20Jun 3, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Toolset for modifying Julia AST and characteristic values☆16Mar 16, 2019Updated 7 years ago
- A command line tool for translation of Flutter ARB files☆16Jul 3, 2026Updated last month
- Structured local memory storage and retrieval for LLM agents☆15May 19, 2026Updated 3 months ago
- Qwen 3.5 in C☆24Mar 28, 2026Updated 4 months ago
- A conversational voice-to-voice open-weights LLM-powered assistant designed to run on high-end consumer or workstation class hardware☆76Jul 6, 2026Updated last month
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 4 months ago
- The popular Matrix effect of falling characters implemented in Flutter☆17Sep 16, 2020Updated 5 years ago
- The Lesma Programming Language☆21Apr 22, 2026Updated 4 months ago
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆73Aug 2, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 11 months ago
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆170Updated this week
- minimal fill-in-the-middle autocomplete for vscode/codium, for use with llama.cpp infill or any openai-compatible server☆23May 10, 2026Updated 3 months ago
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆53Nov 14, 2025Updated 9 months ago
- A desktop GUI for Flux 1.1 Pro built using DelphiFMX For Python☆11Oct 5, 2024Updated last year
- ☆24Jul 15, 2012Updated 14 years ago
- ☆64Updated this week