Collections and tutorials for ROCm
☆32May 25, 2025Updated last year
Alternatives and similar repositories for awesome-rocm
Users that are interested in awesome-rocm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Regex Engine using SIMD and Roaring-Bitmaps☆11Dec 26, 2022Updated 3 years ago
- DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU.☆28Nov 26, 2025Updated 10 months ago
- ☆12Sep 1, 2023Updated 3 years ago
- AI based singing voice synthesis database generator☆13Aug 12, 2022Updated 4 years ago
- 🚀🚀🚀 This repository lists some awesome public CUDA, cuda-python, cuBLAS, cuDNN, CUTLASS, TensorRT, TensorRT-LLM, Triton, TVM, MLIR, PT…☆513Aug 2, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- 气垫船计划——免费、去中心化的北京大学往年题资料库☆13Mar 12, 2018Updated 8 years ago
- PKU Mirror Frontend☆11Apr 5, 2025Updated last year
- Memory access traces of 5 Linux X applications☆11Feb 28, 2021Updated 5 years ago
- The Taichi MPI demos with MPI4Py☆13Nov 3, 2022Updated 3 years ago
- TA's implementation for the project of Computer Architecture and Intelligent Chip Design (23 Spring)☆10May 20, 2023Updated 3 years ago
- ☆16Oct 25, 2024Updated last year
- ☆13Aug 31, 2023Updated 3 years ago
- Open source simulator for porous media flow☆14Oct 15, 2022Updated 3 years ago
- ☆38Nov 28, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Software Analysis and Verification Group☆15Jul 14, 2026Updated 2 months ago
- A Low-Overhead tool for Floating-Point Exception Detection in NVIDIA GPUs☆16Dec 17, 2024Updated last year
- Official PyTorch implementation of "Query-Efficient and Scalable Black-Box Adversarial Attacks on Discrete Sequential Data via Bayesian O…☆26Sep 26, 2023Updated 3 years ago
- OpenCL tool to detect buffer overflows in GPU kernels☆23Jan 7, 2019Updated 7 years ago
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆63Feb 25, 2026Updated 7 months ago
- The AMD rocAL is designed to efficiently decode and process images and videos from a variety of storage formats and modify them through a…☆26Updated this week
- Data Parallel Extension for Numba☆89Sep 26, 2025Updated last year
- A compiler written in Mojo 🔥 and generates RISC-V assembly☆17Jun 25, 2024Updated 2 years ago
- ☆16Feb 11, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- AI Tensor Engine for ROCm☆571Updated this week
- [COLM 2024] SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models☆25Oct 5, 2024Updated last year
- ☆20Sep 19, 2026Updated last week
- HIP backend patch for Numba, the NumPy aware dynamic Python compiler using LLVM.☆22Updated this week
- ☆24Jul 20, 2023Updated 3 years ago
- CPU Memory Compiler and Parallel programing☆26Nov 18, 2024Updated last year
- Fast alignment tool based on bwa-mem☆29Feb 22, 2026Updated 7 months ago
- PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation☆33Nov 16, 2024Updated last year
- themoviedb.org TV Show scraper in Python for Kodi 18 (Leia) or later.☆26Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆35Mar 17, 2026Updated 6 months ago
- A dynamic binary instrumentation tool for tracing and analyzing GPU kernel instructions.☆80Updated this week
- Asset pack for the Vulkan samples repository☆31May 16, 2023Updated 3 years ago
- Accelerated computing with HIP☆29Updated this week
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models☆32Apr 2, 2026Updated 5 months ago
- A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of …☆359Jun 10, 2025Updated last year
- ☆25Jun 24, 2022Updated 4 years ago