parallel algorithm based on cuda
☆60Nov 27, 2017Updated 8 years ago
Alternatives and similar repositories for algorithms-cuda
Users that are interested in algorithms-cuda are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Shared memory overlap-and-save method for NVIDIA GPUs using CUDA☆18Aug 21, 2025Updated last year
- General Industry Camera Driver☆13Dec 22, 2020Updated 5 years ago
- ICML2017 MEC: Memory-efficient Convolution for Deep Neural Network C++实现(非官方)☆17Apr 9, 2019Updated 7 years ago
- Implementation of 3d non-separable convolution using CUDA & FFT Convolution☆20Jan 15, 2019Updated 7 years ago
- A program that runs a sobel filter edge detection algorithm on an image using a single thread on the CPU, another using OpenMP to paralle…☆10Oct 18, 2017Updated 8 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Fast binary matrix product on CPU☆10Feb 11, 2016Updated 10 years ago
- ☆269Jan 14, 2018Updated 8 years ago
- How to plot for papers, slides, demos, etc.☆10Apr 7, 2022Updated 4 years ago
- ☆19Nov 2, 2022Updated 3 years ago
- genetic algorithm in CUDA☆25Dec 4, 2018Updated 7 years ago
- Code for reproducing experiments performed for Accoridon☆13Jun 11, 2021Updated 5 years ago
- Common API (C++ and Python) for Fast Fourier Transform HPC libraries (publish-only mirror)☆11Dec 3, 2025Updated 9 months ago
- Code for the blog post on few-shot classification via task representation and communication.☆18May 24, 2017Updated 9 years ago
- Some source code about matrix multiplication implementation on CUDA☆34Sep 12, 2018Updated 8 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Mixed-Radix DIT FFT in C++11☆13Aug 27, 2018Updated 8 years ago
- Prediction pipeline to generate prognosis predictors for Ebola Virus Disease☆13Feb 22, 2016Updated 10 years ago
- This code generates the filter weights for polyphase filter banks with arbitrary numbers of channels, and with configurable windows.☆31Jul 10, 2024Updated 2 years ago
- Windows Visual Studio Solutions for class "Introduction to Parallel Programming"☆19Nov 17, 2018Updated 7 years ago
- SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling☆16Aug 14, 2025Updated last year
- GHive: Accelerating Analytical Query Processing in Apache Hive via CPU-GPU Heterogeneous Computing.☆15Nov 8, 2023Updated 2 years ago
- Java memory model research☆10Jan 26, 2021Updated 5 years ago
- A Caffe implementation of PSROI-Align☆55Jan 1, 2018Updated 8 years ago
- Case studies constitute a modern interdisciplinary and valuable teaching practice which plays a critical and fundamental role in the deve…☆14Aug 26, 2018Updated 8 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- CUDA Data Parallel Primitives Library☆438Nov 9, 2018Updated 7 years ago
- ☆19Jan 4, 2024Updated 2 years ago
- Kunpeng Tech Blog: https://kunpengcompute.github.io/☆20Jul 8, 2021Updated 5 years ago
- PyTorch code for full quantization of DNN using BCGD☆14Jul 24, 2019Updated 7 years ago
- Convolutional Neural Network of vgg19 model using Cuda to accelerate☆12Jun 11, 2018Updated 8 years ago
- StarPU Runtime system☆16Sep 22, 2010Updated 16 years ago
- Pytorch implementation of Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning (Chen et al)☆27Jun 23, 2018Updated 8 years ago
- Large matrix multiplication in CUDA☆17Oct 20, 2023Updated 2 years ago
- CMake/GoogleTest/TravisCI/Coveralls/CoverityScan/Doxygen☆10Aug 8, 2019Updated 7 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- CACU's Evolution version☆17Feb 21, 2023Updated 3 years ago
- Parallel quicksort algorithms☆16Aug 8, 2010Updated 16 years ago
- demo code for "Principles on Learning New Features for Effective Dense Matching"☆12Nov 1, 2016Updated 9 years ago
- ☆12Feb 14, 2019Updated 7 years ago
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆22Nov 28, 2025Updated 9 months ago
- SharpNTCIP☆11Apr 26, 2014Updated 12 years ago
- ☆11Apr 22, 2020Updated 6 years ago