An Open Source Kepler GPU Assembler
☆23Jan 23, 2017Updated 9 years ago
Alternatives and similar repositories for KeplerAs
Users that are interested in KeplerAs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- assembler for NVIDIA FERMI. Imported from Google Code☆78Mar 22, 2015Updated 11 years ago
- Third party assembler and GEMM library for NVIDIA Kepler GPU☆87Oct 8, 2019Updated 6 years ago
- Assembler for NVIDIA Volta and Turing GPUs☆248Jan 13, 2022Updated 4 years ago
- Experiments evaluating preemption on the NVIDIA Pascal architecture☆16Nov 10, 2016Updated 9 years ago
- ☆17Aug 9, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An unofficial cuda assembler, for all generations of SASS, hopefully :)☆629Apr 20, 2023Updated 3 years ago
- ☆49Dec 11, 2020Updated 5 years ago
- aliyun pk report 2012☆20Oct 31, 2012Updated 13 years ago
- Efficient LDA solution on GPUs.☆24Aug 20, 2018Updated 8 years ago
- Efficient CUDA Stream Compaction Library☆34Jun 9, 2023Updated 3 years ago
- CUDA Tensor Transpose (cuTT) library☆55Aug 10, 2017Updated 9 years ago
- NeuroSync: A Scalable and Accurate Brain Simulation System using Safe and Efficient Speculation (HPCA 2022)☆14Nov 9, 2022Updated 3 years ago
- Use tensor core to calculate back-to-back HGEMM (half-precision general matrix multiplication) with MMA PTX instruction.☆13Nov 3, 2023Updated 2 years ago
- Benchmark of TVM quantized model on CUDA☆112Jun 19, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Flexible GPGPU instrumentation☆91Oct 10, 2019Updated 6 years ago
- Assembler for NVIDIA Maxwell architecture☆1,074Jan 3, 2023Updated 3 years ago
- Multiple 1-stencil implementations using nvidia cuda.☆12Dec 2, 2017Updated 8 years ago
- A Benchmark Toolkit for Assembly Instructions Using the LLVM JIT☆18Oct 26, 2020Updated 5 years ago
- chef cookbook to install Apache Spark☆10Jul 17, 2015Updated 11 years ago
- ☆45Apr 3, 2022Updated 4 years ago
- ☆13Jun 22, 2023Updated 3 years ago
- Check various boost headers impact on the compilation time☆13Jul 11, 2021Updated 5 years ago
- An open-source framework for optimizing binary image processing algorithms.☆16Feb 25, 2021Updated 5 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Utilities and eXtensionS (UXS) library is a collection of useful (template) classes and functions developed upon standard C++ library☆13Updated this week
- GPU implementation of Winograd convolution☆10Oct 23, 2017Updated 8 years ago
- C library plusifier☆11Nov 13, 2021Updated 4 years ago
- Proof of concept prototype to perform distributed training using BVLC/caffe, based on a parameter server implementation using MPI. Data p…☆13May 7, 2015Updated 11 years ago
- ☆48Nov 1, 2025Updated 10 months ago
- Relief Mapping Demo☆13Aug 18, 2011Updated 15 years ago
- This is a demo how to write a high performance convolution run on apple silicon☆56Feb 8, 2022Updated 4 years ago
- Decuda and cudasm, the CUDA binary utilities package. Low-level tools for NVidia G80 GPUs.☆107Jul 24, 2010Updated 16 years ago
- PLCT实验室2019年开放日资料(OpenDay-2019)☆11Dec 20, 2019Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- coffeescript based hardware description language☆14Jan 14, 2022Updated 4 years ago
- ☆10Apr 24, 2023Updated 3 years ago
- A GCC plugin to insert pytest-like assert introspections☆19Jun 6, 2020Updated 6 years ago
- This project is to identify buildings in satellite using Unet and masking method☆11Apr 12, 2026Updated 5 months ago
- Several common methods of matrix multiplication are implemented on CPU and Nvidia GPU using C++11 and CUDA.☆14Feb 8, 2023Updated 3 years ago
- mmdetection -> TVM☆15Aug 22, 2020Updated 6 years ago
- It's a project combined with hardware and software, the goal is to make a smart watch based on esp8266 chip. The smart watch has so many …☆10Jul 9, 2019Updated 7 years ago