CUDA by practice
☆140Jan 7, 2020Updated 6 years ago
Alternatives and similar repositories for CUDA_by_practice
Users that are interested in CUDA_by_practice are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Winograd-based convolution implementation in OpenCL☆29Jan 22, 2017Updated 9 years ago
- Tensorflow Operation Wrapper of cppjieba (Chinese Word Segamentation)☆10Oct 21, 2019Updated 6 years ago
- ☆19Apr 6, 2024Updated 2 years ago
- Sparse matrix computation library for GPU☆59Jul 12, 2020Updated 6 years ago
- Implementation of vDNN++; an improvement over vDNN☆18Dec 7, 2018Updated 7 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Publish sensor data from iOS device to ROS topic☆16Oct 4, 2016Updated 9 years ago
- ☆50Jun 27, 2019Updated 7 years ago
- Sim-to-Real via Sim-to-Sim using fast.ai's U-net☆10Nov 25, 2019Updated 6 years ago
- generative-camouflaged-spam-detector☆11Aug 20, 2020Updated 6 years ago
- Source code of the paper "OpSparse: a Highly Optimized Framework for Sparse General Matrix Multiplication on GPUs"☆16Aug 23, 2022Updated 4 years ago
- An example platform integrating a flask client, a golang server with mongoDb and gRPC for communication☆10Dec 19, 2017Updated 8 years ago
- A visual dataflow programming language for NVIDIA's RAPIDS, based on AlvarBer/Persimmon☆14Jun 1, 2019Updated 7 years ago
- study of cutlass☆22Nov 10, 2024Updated last year
- The C++ implementation of Multi-H algorithm, which is a multi-plane fitting technique. If you use this work for Academic purposes, pleas…☆32Feb 26, 2019Updated 7 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MAchine Micro Management UTilities☆11Nov 5, 2020Updated 5 years ago
- Dynamic matrix type and algorithms for sparse matrices☆24Feb 12, 2025Updated last year
- Effective transpose on Hopper GPU☆29Sep 6, 2025Updated last year
- ☆10Jul 16, 2017Updated 9 years ago
- C++ library to parse WARC files☆11Jan 27, 2019Updated 7 years ago
- OpenGL Panorama Player.☆11May 9, 2018Updated 8 years ago
- It is an annoying thing of preparing the openCL environment, so I wapper the initialization part of OpenCL and setting parameters for ker…☆16May 16, 2018Updated 8 years ago
- Source code that accompanies The CUDA Handbook.☆599Aug 15, 2026Updated last month
- I optimized some code for the European Space Agency, achieving significant speedups - using OpenMP, SSE, Eigen and CUDA. This is the resu…☆16Aug 4, 2017Updated 9 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Convolutional Neural Network of vgg19 model using Cuda to accelerate☆12Jun 11, 2018Updated 8 years ago
- A desktop pet program.☆11Aug 22, 2022Updated 4 years ago
- AutoParBench is a benchmark framework to evaluate compilers and tools designed to automatically insert OpenMP directives.☆12Nov 6, 2020Updated 5 years ago
- A simple example how to use gstreamer-1.0 appsrc and appsink without signals☆27Dec 13, 2016Updated 9 years ago
- ☆12Sep 28, 2023Updated 2 years ago
- C++ heterogeneous and lock-free containers☆13Sep 5, 2018Updated 8 years ago
- Large matrix multiplication in CUDA☆17Oct 20, 2023Updated 2 years ago
- ☆40Feb 28, 2020Updated 6 years ago
- We have developed Symbol Demonstration Direct Preference Optimization (SymDPO) and validating its effectiveness across multiple benchmark…☆23Nov 22, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆13Nov 15, 2017Updated 8 years ago
- Lecture on SIMD units☆11Feb 28, 2017Updated 9 years ago
- Learn CUDA Programming, published by Packt☆1,265Dec 30, 2023Updated 2 years ago
- A flexible, templated GPU library of neighbor search algorithms.☆11Aug 11, 2026Updated last month
- This is the (evolving) reading list for the seminar.☆62Nov 4, 2020Updated 5 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 7 months ago
- An automatic filter-branch of Go libraries from the great Vitess project.☆14Jun 2, 2019Updated 7 years ago