Parallel Computing starter project to build GPU & CPU kernels in CUDA & C++ and call them from Python without a single line of CMake using PyBind11
β31Oct 14, 2025Updated 11 months ago
Alternatives and similar repositories for PyBindToGPUs
Users that are interested in PyBindToGPUs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simple cross-platform speed & memory-efficiency benchmark for the most common hash-table implementations in the C++ worldβ12Dec 9, 2022Updated 3 years ago
- GPU-accelerated Schulze voting method in Python, Numba, CUDA, and Mojo π₯, using ideas from Algebraic Graph Theoryβ19Updated this week
- NetworkX-like Python experience for Postgres, SQLite, MongoDB, and Neo4Jβ32Sep 20, 2026Updated 2 weeks ago
- Optimizing bit-level Jaccard Index and Population Counts for large-scale quantized Vector Search via Harley-Seal CSA and Lookup Tablesβ22May 18, 2025Updated last year
- Thrust, CUB, TBB, AVX2, AVX-512, CUDA, OpenCL, OpenMP, Metal, and Rust - all it takes to sum a lot of numbers fast!β119Jul 22, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Mixed-precision numerics benchmarks in Rust and Python - covering GEMMs, SYRKs, DOTs, and higher-level BLAS and LAPACK-style functionalitβ¦β33Oct 1, 2026Updated last week
- Benchmark suite that compares vector search engines against each other on billion-scale datasets, from in-memory HNSW libraries like USeaβ¦β35Oct 1, 2026Updated last week
- JSON encoder and decoder for python written in C/C++β11Jan 22, 2024Updated 2 years ago
- How to call NVTX from Fortranβ13Jun 25, 2025Updated last year
- This is an advanced tutorial to OpenACC and OpenMP.β15Feb 17, 2022Updated 4 years ago
- Structured Generation Evalsβ14Sep 25, 2024Updated 2 years ago
- Two implementations of ZeRO-1 optimizer sharding in JAXβ14Jun 11, 2023Updated 3 years ago
- Comparing performance-oriented string-processing libraries for substring search, multi-pattern matching, hashing, edit-distances, sketchiβ¦β160Oct 1, 2026Updated last week
- β17Aug 11, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- QR Decoder for Ruby. Uses libdecodeqrβ22Feb 24, 2009Updated 17 years ago
- A Z-order (Morton-code) like coordinate system as template library for arbitrary dimensions.β12Nov 24, 2019Updated 6 years ago
- CLI utilty to work out proper constants for vpternlogic instructionβ13Jan 22, 2023Updated 3 years ago
- A package for defining deep learning models using categorical algebraic expressions.β61Jul 27, 2024Updated 2 years ago
- [ICML 2024] Recurrent Distance Filtering for Graph Representation Learningβ15Jun 10, 2024Updated 2 years ago
- A Lightweight Graph Processing Framework for Multi-GPUsβ14Apr 15, 2015Updated 11 years ago
- Slides library written with AngularJSβ20Apr 30, 2012Updated 14 years ago
- Source Control for People Who Don't Like Source Controlβ25Feb 2, 2010Updated 16 years ago
- Efficient PScan implementation in PyTorchβ17Jan 2, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Cross-platform headers-only C++11 framework for streaming data packets through a graph of data-transforming nodesβ15Aug 8, 2015Updated 11 years ago
- Sources for the Oak Ridge Leadership Computing Facility User Documentationβ67Updated this week
- A set of radiation transport mini-applications used for performance optimization on HPC systems.β29Aug 14, 2018Updated 8 years ago
- QMCPACK miniapp: a simplified real space QMC code for algorithm development, performance portability testing, and computer science experiβ¦β27Jul 24, 2024Updated 2 years ago
- β11Jul 20, 2017Updated 9 years ago
- Low-latency NUMA-aware fork-join thread-pool with zero allocations, syscalls, CAS, or false-sharing on the hot path for C, C++, Rust, andβ¦β377Oct 2, 2026Updated last week
- Basic templates for individual user pages in the public/_www directoryβ11Jun 5, 2026Updated 4 months ago
- Contain some materials about CXL.β20Feb 29, 2024Updated 2 years ago
- Better LaTeX that compiles to LaTeXβ29Sep 27, 2026Updated last week
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Finite state transducer library. Minimalistic pure C implementation.β15Apr 23, 2020Updated 6 years ago
- Focused on fast experimentation and simplicityβ77Dec 24, 2024Updated last year
- The C++ library - "glib" built for ios-simulator and ios-devicesβ11Dec 24, 2014Updated 11 years ago
- β28Sep 28, 2026Updated last week
- A repository to unravel the language of GPUs, making their kernel conversations easy to understandβ216Jun 1, 2025Updated last year
- Parallel implementation of the Advanced Encryption Standard.β11Nov 13, 2018Updated 7 years ago
- β94Jul 5, 2024Updated 2 years ago