π GPU load-balancing library for regular and irregular computations.
β67Jun 25, 2026Updated 2 months ago
Alternatives and similar repositories for loops
Users that are interested in loops are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β14Apr 24, 2024Updated 2 years ago
- mini is miniβ20Jan 19, 2020Updated 6 years ago
- Runs a single CUDA/OpenCL kernel, taking its source from a file and arguments from the command-lineβ26Updated this week
- Open-source library for Graph Streaming. Solves the connected components problem using sub-linear space. Published in SIGMOD'22.β11Apr 6, 2026Updated 5 months ago
- A vectorizable multi-dimensional iterator for C++ using the Coroutines TSβ12Jun 5, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- CUDA Dynamic Memory Allocator for SOA Data Layoutβ40Dec 29, 2021Updated 4 years ago
- Source code for the paper: Accelerating Dynamic Graph Analytics on GPUsβ30Jun 19, 2023Updated 3 years ago
- β20Jan 17, 2024Updated 2 years ago
- β667Updated this week
- cuASR: CUDA Algebra for Semiringsβ52Aug 22, 2022Updated 4 years ago
- Chapel HyperGraph Library (CHGL) - HPC-class Hypergraphs in Chapelβ35Oct 29, 2020Updated 5 years ago
- Generate simple index ranges in C++ and CUDA C++β39Jun 14, 2023Updated 3 years ago
- Efficient and High-quality Graph Coloring on the GPUβ16Apr 3, 2022Updated 4 years ago
- β11Apr 10, 2019Updated 7 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Source code supporting the High Performance Graphics 2022 paper: Supporting Unified Shader Specialization by Co-opting C++ Featuresβ14Jul 9, 2022Updated 4 years ago
- β18Oct 15, 2020Updated 5 years ago
- Record GPU memory accesses of a CUDA program and visualize the access pattern in a browserβ13Nov 17, 2020Updated 5 years ago
- Multi-GPU dynamic scheduler using PGAS style cross-GPU communicationβ29Jul 23, 2023Updated 3 years ago
- Programmable CUDA/C++ GPU Graph Analyticsβ1,098Feb 28, 2026Updated 6 months ago
- Evaluating different memory managers for dynamic GPU memoryβ26Dec 16, 2020Updated 5 years ago
- GPUDirect Async implementation of HPGMG-FV CUDAβ11May 11, 2018Updated 8 years ago
- Department of Energy Standard Utility Libraryβ34Updated this week
- A pattern-based algorithmic autotuner for graph processing on GPUs.β33Jun 25, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A Collection of Parallel Algorithms for Computational Geometryβ12Mar 10, 2022Updated 4 years ago
- Matrix multiplication on GPUs for matrices stored on a CPU. Similar to cublasXt, but ported to both NVIDIA and AMD GPUs.β33Apr 2, 2025Updated last year
- Fast SGEMM emulation on Tensor Coresβ17Feb 16, 2025Updated last year
- Statistics on GPUsβ33May 5, 2026Updated 4 months ago
- Artifact for PPoPP20 "Understanding and Bridging the Gaps in Current GNN Performance Optimizations"β42Nov 16, 2021Updated 4 years ago
- Using C++ magic to capture CUDA kernels and tune them with Kernel Tunerβ22Aug 25, 2026Updated 3 weeks ago
- A C++based implementation of the TeaLeaf heat conduction mini-app. This implementation of TeaLeaf replicates the functionality of the refβ¦β25Aug 11, 2024Updated 2 years ago
- GenDP: A Dynamic Programming Framework for Genome Sequencing Analysisβ17Jan 12, 2024Updated 2 years ago
- A reference implementation of std::simd, providing data parallel types in the C++ standardβ14Mar 9, 2020Updated 6 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Artifact for OSDI'21 GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUs.β71Mar 2, 2023Updated 3 years ago
- SparseTIR: Sparse Tensor Compiler for Deep Learningβ145Mar 31, 2023Updated 3 years ago
- Scale-out system monitoringβ27Updated this week
- Heron: Automatically Constrained High-Performance Library Generation for Deep Learning Acceleratorsβ24Jan 30, 2024Updated 2 years ago
- Repository for artifact evaluation of ASPLOS 2023 paper "SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning"β25Feb 24, 2023Updated 3 years ago
- β31Aug 28, 2020Updated 6 years ago
- A language and compiler for irregular tensor programs.β153Updated this week