π GPU load-balancing library for regular and irregular computations.
β67Jun 25, 2026Updated 2 months ago
Alternatives and similar repositories for loops
Users that are interested in loops are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β€οΈ CUDA/C++ GPU graph analytics simplified.β32Sep 19, 2022Updated 3 years ago
- β14Apr 24, 2024Updated 2 years ago
- mini is miniβ20Jan 19, 2020Updated 6 years ago
- Runs a single CUDA/OpenCL kernel, taking its source from a file and arguments from the command-lineβ26Jun 10, 2026Updated 2 months ago
- Open-source library for Graph Streaming. Solves the connected components problem using sub-linear space. Published in SIGMOD'22.β11Apr 6, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A vectorizable multi-dimensional iterator for C++ using the Coroutines TSβ12Jun 5, 2022Updated 4 years ago
- CUDA Dynamic Memory Allocator for SOA Data Layoutβ40Dec 29, 2021Updated 4 years ago
- Source code for the paper: Accelerating Dynamic Graph Analytics on GPUsβ30Jun 19, 2023Updated 3 years ago
- β20Jan 17, 2024Updated 2 years ago
- β661Updated this week
- cuASR: CUDA Algebra for Semiringsβ51Aug 22, 2022Updated 4 years ago
- Generate simple index ranges in C++ and CUDA C++β39Jun 14, 2023Updated 3 years ago
- β11Aug 8, 2021Updated 5 years ago
- LonestarGPU: Irregular algorithms parallelized for GPUsβ38Nov 11, 2019Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β11Apr 10, 2019Updated 7 years ago
- Source code supporting the High Performance Graphics 2022 paper: Supporting Unified Shader Specialization by Co-opting C++ Featuresβ14Jul 9, 2022Updated 4 years ago
- β18Oct 15, 2020Updated 5 years ago
- Record GPU memory accesses of a CUDA program and visualize the access pattern in a browserβ13Nov 17, 2020Updated 5 years ago
- Multi-GPU dynamic scheduler using PGAS style cross-GPU communicationβ29Jul 23, 2023Updated 3 years ago
- Programmable CUDA/C++ GPU Graph Analyticsβ1,095Feb 28, 2026Updated 6 months ago
- GPUDirect Async implementation of HPGMG-FV CUDAβ11May 11, 2018Updated 8 years ago
- PilotFish harvests the free GPU cycles of cloud gaming with deep learning trainingβ14Jul 2, 2022Updated 4 years ago
- Department of Energy Standard Utility Libraryβ34Aug 25, 2026Updated last week
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A pattern-based algorithmic autotuner for graph processing on GPUs.β33Jun 25, 2025Updated last year
- A Collection of Parallel Algorithms for Computational Geometryβ12Mar 10, 2022Updated 4 years ago
- Matrix multiplication on GPUs for matrices stored on a CPU. Similar to cublasXt, but ported to both NVIDIA and AMD GPUs.β33Apr 2, 2025Updated last year
- Fast SGEMM emulation on Tensor Coresβ17Feb 16, 2025Updated last year
- Statistics on GPUsβ33May 5, 2026Updated 3 months ago
- Artifact for PPoPP20 "Understanding and Bridging the Gaps in Current GNN Performance Optimizations"β42Nov 16, 2021Updated 4 years ago
- Using C++ magic to capture CUDA kernels and tune them with Kernel Tunerβ22Aug 25, 2026Updated last week
- A C++based implementation of the TeaLeaf heat conduction mini-app. This implementation of TeaLeaf replicates the functionality of the refβ¦β25Aug 11, 2024Updated 2 years ago
- GenDP: A Dynamic Programming Framework for Genome Sequencing Analysisβ17Jan 12, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A reference implementation of std::simd, providing data parallel types in the C++ standardβ14Mar 9, 2020Updated 6 years ago
- Artifact for OSDI'21 GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUs.β71Mar 2, 2023Updated 3 years ago
- SparseTIR: Sparse Tensor Compiler for Deep Learningβ145Mar 31, 2023Updated 3 years ago
- Heron: Automatically Constrained High-Performance Library Generation for Deep Learning Acceleratorsβ24Jan 30, 2024Updated 2 years ago
- Scale-out system monitoringβ26Updated this week
- β31Aug 28, 2020Updated 6 years ago
- A language and compiler for irregular tensor programs.β152Aug 16, 2026Updated 2 weeks ago