Thrust, CUB, TBB, AVX2, AVX-512, CUDA, OpenCL, OpenMP, Metal, and Rust - all it takes to sum a lot of numbers fast!
β119Jul 22, 2025Updated 11 months ago
Alternatives and similar repositories for ParallelReductionsBenchmark
Users that are interested in ParallelReductionsBenchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- GPU-accelerated Schulze voting method in Python, Numba, CUDA, and Mojo π₯, using ideas from Algebraic Graph Theoryβ19Oct 28, 2025Updated 8 months ago
- Parallel Computing starter project to build GPU & CPU kernels in CUDA & C++ and call them from Python without a single line of CMake usinβ¦β31Oct 14, 2025Updated 9 months ago
- Link to this library and it will log all the LibC functions you are calling and how much time you are spending in them!β22Jan 4, 2025Updated last year
- A Halide backend for ONNXβ12Nov 5, 2019Updated 6 years ago
- Optimizing bit-level Jaccard Index and Population Counts for large-scale quantized Vector Search via Harley-Seal CSA and Lookup Tablesβ22May 18, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- GPGPU array on Vulkanβ17Jun 3, 2023Updated 3 years ago
- Benchmark suite that compares vector search engines against each other on billion-scale datasets, from in-memory HNSW libraries like USeaβ¦β35May 2, 2026Updated 2 months ago
- Tiny Semantic Versioning (SemVer) library with LLMs and GitHub CI, that doesn't depend on 300K lines of JavaScript code and fits in a sinβ¦β28Updated this week
- Case Studies for Halide performance against C++ and OpenCLβ36Oct 9, 2013Updated 12 years ago
- Concurrent CPU-GPU Programming using Task Modelsβ110Dec 19, 2019Updated 6 years ago
- Collection of samples and utilities for using ComputeCpp, Codeplay's SYCL implementationβ326Aug 11, 2023Updated 2 years ago
- β11Jul 13, 2022Updated 4 years ago
- Apache Arrow-compatible space-efficient "tape" class in pure Rust to be used with StringZilla for GPU, NUMA, and disk transfers of variabβ¦β31Updated this week
- A wrapper around Python's ctypes for Nim-specific function signatures.β12Dec 12, 2017Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- a Halide language To MLIR compiler.β26Aug 30, 2021Updated 4 years ago
- C++ code for RingQueue, SharedMemory and Semaphoreβ20Aug 29, 2020Updated 5 years ago
- Awesome implicit data structuresβ24Oct 6, 2019Updated 6 years ago
- A community-oriented list of useful NUMA-related libraries, tools, and other resourcesβ78May 12, 2026Updated 2 months ago
- My very own vxsort re-implemented with "modern" C++ by a complete idiot (in C++)β33Jun 12, 2026Updated last month
- book for Halide language programmingβ13Sep 8, 2021Updated 4 years ago
- OpenCl block for Cinderβ14Feb 23, 2016Updated 10 years ago
- study of cutlassβ22Nov 10, 2024Updated last year
- CuPBoP-AMD is a CUDA translator that translates CUDA programs at NVVM IR level to HIP-compatible IR that can run on AMD GPUs.β42Nov 19, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Multi-GPU Framework for Voxel Grid Computationsβ69Jul 7, 2026Updated last week
- CNNs in Halideβ22Oct 22, 2015Updated 10 years ago
- outline and links for PLDI 2022 tutorialβ17Jun 13, 2022Updated 4 years ago
- Experimental ranges for CUDAβ25Feb 1, 2019Updated 7 years ago
- The project provides high-performance concurrency, enabling highly parallel computation.β283Jun 28, 2026Updated 3 weeks ago
- Compiler for multiple programming models (SYCL, C++ standard parallelism, HIP/CUDA) for CPUs and GPUs from all vendors: The independent, β¦β1,908Updated this week
- Dataframe for Nimβ15Jul 25, 2019Updated 6 years ago
- KNoC is a Kubernetes Virtual Kubelet that uses an HPC cluster as the container execution environmentβ21Feb 1, 2023Updated 3 years ago
- The translator that supports translating NVPTX to SPIR-V. This translator is modified from LLVM-SPIR-V Translator.β45Oct 25, 2021Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A streamlined CMake build system foundation for developing HPC softwareβ293Updated this week
- A simple profiler to count Nvidia PTX assembly instructions of OpenCL/SYCL/CUDA kernels for roofline model analysis.β59Mar 20, 2025Updated last year
- Distributed Performance-portable Stencil Compuitationβ10Jul 9, 2023Updated 3 years ago
- An efficient C++20 GPU numerical computing library with Python-like syntaxβ1,438Updated this week
- An expression template based linear algebra library running completely on the GPU using CUDAβ26Jun 24, 2021Updated 5 years ago
- Universal adapter to transform class methodsβ15Oct 17, 2017Updated 8 years ago
- Open Source Parallel STL implementationβ531Jan 26, 2024Updated 2 years ago