implementation of floating-point radix sorting based on CUDA
☆33Feb 10, 2020Updated 6 years ago
Alternatives and similar repositories for CUDA_radix_sort
Users that are interested in CUDA_radix_sort are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CUDA implementation of parallel radix sort using Blelloch scan☆70Feb 29, 2024Updated 2 years ago
- CUDA PTX-ISA Document 中文翻译版☆58Sep 29, 2025Updated 11 months ago
- ☆14Oct 4, 2018Updated 7 years ago
- 🐱 ncnn int8 模型量化评估☆14Oct 10, 2022Updated 3 years ago
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A Triton JIT runtime and ffi provider in C++☆40Aug 7, 2026Updated 3 weeks ago
- ☆14Apr 24, 2024Updated 2 years ago
- Screen space global illumination for interactive mixed reality☆14Dec 13, 2017Updated 8 years ago
- Pytorch Implementation of Signed Neuron with Memory: Towards Simple, Accurate and High-Efficient ANN-SNN Conversion, IJCAI 2022☆22Dec 14, 2022Updated 3 years ago
- ☆73Jan 6, 2025Updated last year
- A .NET / C# wrapper for the Mikktspace tangent generation algorithm☆14Jul 11, 2021Updated 5 years ago
- A .NET wrapper for BinomialLLC's Basis Universal Supercompressed GPU Texture Codec☆10Nov 22, 2022Updated 3 years ago
- Unity ScriptedImporter for Unreal Datasmith bundles☆12Dec 6, 2020Updated 5 years ago
- A Winograd Minimal Filter Implementation in CUDA☆31Aug 25, 2021Updated 5 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆13Jan 23, 2021Updated 5 years ago
- CUDA project for uni subject☆26Oct 26, 2020Updated 5 years ago
- Reader/Writer for Instanced 3DTiles (i3dm)☆17Jul 24, 2024Updated 2 years ago
- Parallel Prefix Sum (Scan) with CUDA☆30Jun 22, 2024Updated 2 years ago
- ENet-caffe uses TensorRT to speed up☆10Apr 25, 2019Updated 7 years ago
- Attempting to implement VXGI for NCCA Masterclass assignment☆15Mar 21, 2018Updated 8 years ago
- A GPU radix sorter using the OpenGL compute shader.☆20Dec 24, 2015Updated 10 years ago
- ☆17Feb 25, 2026Updated 6 months ago
- Awesome MoE Diffusion Models☆21Mar 25, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆10Sep 27, 2023Updated 2 years ago
- Standalone Flash Attention v2 kernel without libtorch dependency☆113Sep 10, 2024Updated last year
- a .NET compatible P/Invoke wrapper around meshoptimizer☆27Jun 10, 2019Updated 7 years ago
- Sparse kernels for GNNs based on TVM☆17Nov 18, 2020Updated 5 years ago
- ☆16Mar 13, 2018Updated 8 years ago
- End to end Tensor IR/DSL stack for deploying deep learning workloads to hardwares☆10Oct 25, 2021Updated 4 years ago
- Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruct…☆568Sep 8, 2024Updated last year
- A polyhedral mesh generation library for Unity based on the halfedge data structure☆42Aug 10, 2026Updated 3 weeks ago
- This is the implementation for paper: AdaTune: Adaptive Tensor Program CompilationMade Efficient (NeurIPS 2020).☆14May 16, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- API serving for your diffusers models☆11Jan 19, 2024Updated 2 years ago
- Source code for the evaluated benchmarks and proposed cache management technique, GRASP, in [Faldu et al., HPCA'20].☆18Jan 23, 2020Updated 6 years ago
- GLTF 2.0 Typescript Interface Generator☆25Sep 9, 2021Updated 4 years ago
- diffusers with search engine☆12Aug 12, 2026Updated 3 weeks ago
- [DATE'23] The official code for paper <CLAP: Locality Aware and Parallel Triangle Counting with Content Addressable Memory>☆24May 25, 2026Updated 3 months ago
- ☆62Jul 3, 2025Updated last year
- Code for GridNet☆18Sep 8, 2017Updated 8 years ago