☆24May 12, 2026Updated 4 months ago
Alternatives and similar repositories for llamacpp-gfx-906-turbo
Users that are interested in llamacpp-gfx-906-turbo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆87Jun 23, 2026Updated 2 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆37Aug 25, 2026Updated 3 weeks ago
- llama.cpp-gfx906☆143Aug 23, 2026Updated 3 weeks ago
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆17Feb 21, 2026Updated 6 months ago
- Random AI notes for working with local models or playing around with random machine learning bits.☆63Jun 7, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 9 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆435Feb 20, 2026Updated 6 months ago
- Collections of multimodal search libraries, service and research papers☆19Apr 18, 2025Updated last year
- Proxy for OpenAI☆16Sep 2, 2025Updated last year
- ☆20Aug 19, 2025Updated last year
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆68Sep 5, 2026Updated last week
- A llamacpp wrapper to manage and monitor your llama server instance over a web ui.☆27Jun 16, 2026Updated 2 months ago
- A lightweight chat interface for interacting with local models, featuring persistent memory using a seamless SQLite database to store you…☆34Sep 15, 2025Updated last year
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆70May 4, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A repository to store helpful information and emerging insights in regard to LLMs☆21Oct 27, 2023Updated 2 years ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- RDNA-native LLM inference engine in Rust.☆626Updated this week
- Gives agents a real browser. URL in, pruned snapshot out. Replaces Playwright, Selenium, Puppeteer. Zero deps, zero wasted tokens.☆46Aug 31, 2026Updated 2 weeks ago
- Loader extension for tabbyAPI in SillyTavern☆27Jun 30, 2025Updated last year
- One-click LLM server with TurboQuant Llama CPP engine☆18Apr 16, 2026Updated 4 months ago
- Qwen 3.5 in C☆24Mar 28, 2026Updated 5 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 5 months ago
- ☆29Jul 24, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆32Updated this week
- Polymer Material Design in GWT☆15Feb 26, 2015Updated 11 years ago
- Batch processor to enable large content be digested by Ollama, focused around book processing and translations by default, fully configur…☆36Aug 23, 2026Updated 3 weeks ago
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆75Aug 2, 2026Updated last month
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated last year
- Script Execution service☆13Nov 21, 2016Updated 9 years ago
- Rust CLI for RBMEM (.rbmem), a structured Rust-Brain memory format with timestamp-protected sections, hierarchy, graph relations, Hermes …☆29Jun 14, 2026Updated 3 months ago
- minimal fill-in-the-middle autocomplete for vscode/codium, for use with llama.cpp infill or any openai-compatible server☆24May 10, 2026Updated 4 months ago
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆175Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆53Nov 14, 2025Updated 10 months ago
- Awesome-RAG: a curated list of Retrieval-Augmented Generation☆54Dec 31, 2024Updated last year
- A desktop GUI for Flux 1.1 Pro built using DelphiFMX For Python☆11Oct 5, 2024Updated last year
- ☆19Jul 2, 2025Updated last year
- Simple agent framework using Ollama tool calling☆10Aug 27, 2024Updated 2 years ago
- GFPGAN face reconstruction with ncnn on a bare Raspberry Pi☆14Jan 4, 2023Updated 3 years ago
- The Lily programming language ⚜☆11Aug 12, 2026Updated last month