☆99May 30, 2026Updated 2 months ago
Alternatives and similar repositories for GPU_Programming
Users that are interested in GPU_Programming are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Step by step implementation of a fast softmax kernel in CUDA☆71Jan 6, 2025Updated last year
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 6 months ago
- torch.compile artifacts for common deep learning models, can be used as a learning resource for torch.compile☆19Dec 22, 2023Updated 2 years ago
- ☆14Dec 22, 2024Updated last year
- Companion code for Grokking Megakernels: fuse an entire LLM forward pass into a single CUDA kernel☆23Feb 9, 2026Updated 6 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository documents my 100-day journey of learning and writing CUDA kernels.☆35Mar 29, 2026Updated 4 months ago
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- Our Clone of Orca used for experimentation☆20Oct 15, 2024Updated last year
- My study notes and hands-on projects for CUDA-based GPU programming☆13Dec 11, 2025Updated 7 months ago
- RAPIDS Deployment Documentation☆15Updated this week
- Learn CUDA with PyTorch☆364Jun 1, 2026Updated 2 months ago
- CUDA Learning guide☆570Jun 20, 2024Updated 2 years ago
- Personal solutions to the Triton Puzzles☆21Jul 18, 2024Updated 2 years ago
- BFloat16 Fused Adam Operator for PyTorch☆20Nov 16, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Flash Attention in raw Cuda C beating PyTorch☆39May 14, 2024Updated 2 years ago
- IBM Spectrum LSF - IBM Cloud☆16Sep 30, 2024Updated last year
- ☆19Jun 6, 2025Updated last year
- Hand-Rolled GPU communications library☆96Nov 25, 2025Updated 8 months ago
- Fast CUDA matrix multiplication from scratch☆1,277Sep 2, 2025Updated 11 months ago
- ☆15Feb 13, 2018Updated 8 years ago
- Comparing Deep Learning Inference of Pytorch models running on CPU, CUDA and TensorRT☆17Feb 20, 2022Updated 4 years ago
- Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O☆593Sep 13, 2025Updated 10 months ago
- Fastest kernels written from scratch☆590Sep 18, 2025Updated 10 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Small scale distributed training of sequential deep learning models, built on Numpy and MPI.☆165Oct 19, 2023Updated 2 years ago
- Repository for the CUDA H100 Course☆68Apr 12, 2026Updated 3 months ago
- Setup Cuda☆29May 23, 2024Updated 2 years ago
- A library of replicated state machine algorithms is based on Viewstamped Replication Revisited☆14Feb 6, 2021Updated 5 years ago
- A 120-day CUDA learning plan covering daily concepts, exercises, pitfalls, and references (including “Programming Massively Parallel Proc…☆945Mar 29, 2025Updated last year
- CPU/GPU Implicit & Explicit Finite Element Solver for Large Strains☆25Feb 20, 2026Updated 5 months ago
- A curated list of resources for learning and exploring Triton, OpenAI's programming language for writing efficient GPU code.☆497Mar 10, 2025Updated last year
- Hugging Face Download (Cache) Manager☆22Aug 7, 2022Updated 4 years ago
- GPU Kernels☆225Apr 27, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆49Apr 15, 2024Updated 2 years ago
- 100 Days Of GPU Programming.☆45Nov 7, 2025Updated 9 months ago
- 3D acoustic wave propagation in homogeneous isotropic media using PETSc and Krylov space method☆16Oct 29, 2017Updated 8 years ago
- ☆40Oct 21, 2025Updated 9 months ago
- Workshop of Melown 3D stack☆12Apr 25, 2019Updated 7 years ago
- High-Performance FP32 GEMM on CUDA devices☆126Jan 21, 2025Updated last year
- GiftHub | Analyze social media for gift recommendations☆14Dec 18, 2017Updated 8 years ago