Thrust, CUB, TBB, AVX2, AVX-512, CUDA, OpenCL, OpenMP, Metal, and Rust - all it takes to sum a lot of numbers fast!
β119Jul 22, 2025Updated last year
Alternatives and similar repositories for ParallelReductionsBenchmark
Users that are interested in ParallelReductionsBenchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- GPU-accelerated Schulze voting method in Python, Numba, CUDA, and Mojo π₯, using ideas from Algebraic Graph Theoryβ19Aug 24, 2026Updated 3 weeks ago
- A simple cross-platform speed & memory-efficiency benchmark for the most common hash-table implementations in the C++ worldβ12Dec 9, 2022Updated 3 years ago
- Parallel Computing starter project to build GPU & CPU kernels in CUDA & C++ and call them from Python without a single line of CMake usinβ¦β31Oct 14, 2025Updated 11 months ago
- A Halide backend for ONNXβ12Nov 5, 2019Updated 6 years ago
- WIP Β· CUDA compatibility for Blaze Β· https://bitbucket.org/blaze-lib/blazeβ21Nov 18, 2019Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Optimizing bit-level Jaccard Index and Population Counts for large-scale quantized Vector Search via Harley-Seal CSA and Lookup Tablesβ22May 18, 2025Updated last year
- GPGPU array on Vulkanβ17Jun 3, 2023Updated 3 years ago
- Benchmark suite that compares vector search engines against each other on billion-scale datasets, from in-memory HNSW libraries like USeaβ¦β35Jul 25, 2026Updated last month
- Tiny Semantic Versioning (SemVer) library with LLMs and GitHub CI, that doesn't depend on 300K lines of JavaScript code and fits in a sinβ¦β29Jul 21, 2026Updated last month
- Runs a single CUDA/OpenCL kernel, taking its source from a file and arguments from the command-lineβ26Updated this week
- Case Studies for Halide performance against C++ and OpenCLβ36Oct 9, 2013Updated 12 years ago
- My notes on various HPC papers.β28Jan 7, 2023Updated 3 years ago
- Concurrent CPU-GPU Programming using Task Modelsβ110Dec 19, 2019Updated 6 years ago
- Collection of samples and utilities for using ComputeCpp, Codeplay's SYCL implementationβ326Aug 11, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- β11Jul 13, 2022Updated 4 years ago
- ILP SAT Detailed Routerβ13Apr 14, 2020Updated 6 years ago
- Apache Arrow-compatible space-efficient "tape" class in pure Rust to be used with StringZilla for GPU, NUMA, and disk transfers of variabβ¦β33Jul 31, 2026Updated last month
- a Halide language To MLIR compiler.β26Aug 30, 2021Updated 5 years ago
- C++ code for RingQueue, SharedMemory and Semaphoreβ20Aug 29, 2020Updated 6 years ago
- A Doxygen plugin for MkDocsβ18Dec 4, 2020Updated 5 years ago
- Awesome implicit data structuresβ24Oct 6, 2019Updated 6 years ago
- A community-oriented list of useful NUMA-related libraries, tools, and other resourcesβ78May 12, 2026Updated 4 months ago
- My very own vxsort re-implemented with "modern" C++ by a complete idiot (in C++)β33Jun 12, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- book for Halide language programmingβ13Sep 8, 2021Updated 5 years ago
- OpenCl block for Cinderβ14Feb 23, 2016Updated 10 years ago
- study of cutlassβ22Nov 10, 2024Updated last year
- CuPBoP-AMD is a CUDA translator that translates CUDA programs at NVVM IR level to HIP-compatible IR that can run on AMD GPUs.β41Nov 19, 2023Updated 2 years ago
- Multi-GPU Framework for Voxel Grid Computationsβ69Updated this week
- CNNs in Halideβ22Oct 22, 2015Updated 10 years ago
- outline and links for PLDI 2022 tutorialβ17Jun 13, 2022Updated 4 years ago
- Experimental ranges for CUDAβ25Feb 1, 2019Updated 7 years ago
- Dataframe for Nimβ15Jul 25, 2019Updated 7 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- STXXL: Standard Template Library for Extra Large Data Setsβ498Dec 22, 2023Updated 2 years ago
- Compiler for multiple programming models (SYCL, C++ standard parallelism, HIP/CUDA) for CPUs and GPUs from all vendors: The independent, β¦β1,939Updated this week
- KNoC is a Kubernetes Virtual Kubelet that uses an HPC cluster as the container execution environmentβ21Feb 1, 2023Updated 3 years ago
- Semantic Search demo featuring UForm, USearch, UCall, and StreamLit, to visual and retrieve from image datasets, similar to "CLIP Retrievβ¦β56Dec 29, 2023Updated 2 years ago
- The project provides high-performance concurrency, enabling highly parallel computation.β305Sep 3, 2026Updated 2 weeks ago
- A streamlined CMake build system foundation for developing HPC softwareβ296Updated this week
- A simple profiler to count Nvidia PTX assembly instructions of OpenCL/SYCL/CUDA kernels for roofline model analysis.β59Mar 20, 2025Updated last year