FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs
☆70May 4, 2025Updated last year
Alternatives and similar repositories for vllm-rocm
Users that are interested in vllm-rocm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Triton for AMD MI25/50/60. Development repository for the Triton language and compiler☆35Dec 15, 2025Updated 8 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆434Feb 20, 2026Updated 6 months ago
- triton3.2.0添加mi25/mi50/mi60支持☆14Apr 26, 2025Updated last year
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 8 months ago
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆335Aug 30, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆444Apr 4, 2025Updated last year
- ☆24May 12, 2026Updated 3 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆124Updated this week
- forked from vllm-project/flash-attention☆64May 9, 2026Updated 3 months ago
- ☆103Dec 29, 2020Updated 5 years ago
- Run a 35B MoE model at 10+ tok/s on a $600 Mac mini. Pure C/Metal inference engine streaming experts from SSD on Apple Silicon☆19Apr 20, 2026Updated 4 months ago
- Code and data for Teddy https://arxiv.org/abs/2001.05171.☆15Jun 21, 2022Updated 4 years ago
- Transform your pdfs into anki flashcard with gpt☆12Jun 10, 2024Updated 2 years ago
- Examples, tutorials, and other development resources for the ISPU, an ultralow-power programmable core embedded in STMicroelectronics MEM…☆23Aug 28, 2026Updated last week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Kit's open-sourced agent workspace 🦊 - SOUL.md, memory, knowledge, scripts☆20Apr 4, 2026Updated 5 months ago
- A full-stack document management and AI chat application that enables users to upload, manage, and chat with their documents using AI. Bu…☆16Aug 10, 2025Updated last year
- LLM inference in C/C++☆30Updated this week
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆925Updated this week
- Tries to UI development. Clone of https://www.perplexity.ai/☆11Sep 30, 2023Updated 2 years ago
- Batch for Excel 2013 SPREADSHEETCOMPARE tool☆14Jun 25, 2014Updated 12 years ago
- Gluetun-Sync is an open-source utility designed to automatically update your services when the NAT port changes in Gluetun VPN☆12Sep 25, 2024Updated last year
- ☆14Sep 4, 2024Updated 2 years ago
- Enhancing LangChain prompts to work better with RWKV models☆32May 30, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- Fused BF16 Huffman GEMV Inference kernel☆22Apr 22, 2026Updated 4 months ago
- A library and CLI utilities for managing performance states of NVIDIA GPUs.☆38Oct 6, 2024Updated last year
- ☆59Jul 16, 2026Updated last month
- AI plays Doom — pit Vision Language Models against demons and each other. Solo scenarios, deathmatch arena, 1-4 agents with any OpenAI-co…☆20Mar 12, 2026Updated 5 months ago
- Code scanner to check for issues in prompts and LLM calls☆78Apr 6, 2025Updated last year
- FastAPI WebSocket server for the OpenVoice text-to-speech model.☆12Jun 6, 2024Updated 2 years ago
- LM inference server implementation based on *.cpp.☆292Nov 24, 2025Updated 9 months ago
- Antlr v4 demo 使用Antlr基于AST(抽象语法树)开发了一套DSL,用来解析游戏策划配置的技能。 支持变量定义,函数定义,控制结构。☆18Feb 27, 2018Updated 8 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Structured local memory storage and retrieval for LLM agents☆15May 19, 2026Updated 3 months ago
- RAIL wireless applications. Go to https://github.com/SiliconLabs/application_examples☆13Jul 24, 2026Updated last month
- ☆17Oct 15, 2023Updated 2 years ago
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆31Jul 7, 2026Updated 2 months ago
- ☆29Jul 24, 2026Updated last month
- Prometheus exporter for Linux based GDDR6/GDDR6X VRAM and GPU Core Hot spot temperature reader for NVIDIA RTX 3000/4000 series GPUs.☆26Oct 2, 2024Updated last year
- CI scripts designed to build a Pascal-compatible version of vLLM.☆13Aug 10, 2024Updated 2 years ago