Code for testing the native float16 matrix multiplication performance on Tesla P100 and V100 GPU based on cublasHgemm
☆35Aug 20, 2019Updated 7 years ago
Alternatives and similar repositories for cublasHgemm-P100
Users that are interested in cublasHgemm-P100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tensorflow model export from Python to C++ and inference without using TF library☆17Mar 13, 2019Updated 7 years ago
- C++ CPU inference library for Tensorflow object detection models based on the lightweight Tensorflow C-API.☆15Jun 26, 2018Updated 8 years ago
- DELTA-pytorch:DELTA: Dynamically Optimizing GPU Memory beyond Tensor Recomputation☆12Apr 16, 2024Updated 2 years ago
- ☆10Aug 18, 2016Updated 10 years ago
- outline and links for PLDI 2022 tutorial☆17Jun 13, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- HCC Sample Applications☆13Jan 3, 2017Updated 9 years ago
- [MLSys 2023] Pre-train and Search: Efficient Embedding Table Sharding with Pre-trained Neural Cost Models☆16May 5, 2023Updated 3 years ago
- A Top-Down Profiler for GPU Applications☆24Feb 29, 2024Updated 2 years ago
- A NetWork Generate Names, Based On Conditional RNN, Set Condition And Generate Different Names.☆12May 15, 2017Updated 9 years ago
- ☆14Jan 12, 2022Updated 4 years ago
- TensorRT Int8 Python version sample. TensorRT Int8 Python 实现例子。TensorRT Int8 Pythonの例です☆14Jan 28, 2019Updated 7 years ago
- RDMA Optimization on MXNet☆14Nov 12, 2017Updated 8 years ago
- A GPU particle demo with fluid motion simulated through curl noise.☆15Jun 21, 2012Updated 14 years ago
- Sparse Matrix-Matrix Multiplication Benchmark on Intel Xeon and Xeon Phi (KNC, KNL) from blog post:☆12Sep 25, 2016Updated 9 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Caffe: a fast open framework for deep learning.☆14Jun 2, 2016Updated 10 years ago
- Distributed RPC framework for enterprise SOA infrastructure☆19Mar 22, 2017Updated 9 years ago
- YOLOv3-training-prune☆58Mar 9, 2021Updated 5 years ago
- ☆28Nov 6, 2024Updated last year
- An implementation of our CVPR 2018 work 'Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning'☆43Jul 12, 2019Updated 7 years ago
- PyTorch 1.5 C++ frontend API☆20May 5, 2020Updated 6 years ago
- Code repo for the paper "SpinQuant LLM quantization with learned rotations"☆15Mar 20, 2025Updated last year
- Accelerating Database Operations on a GPU with CUDA☆18Oct 5, 2015Updated 10 years ago
- a plugin for QtCreator IDE to easily create new Tufão projects☆13Oct 7, 2015Updated 10 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- MobileSAM のエンコーダー/デコーダーをONNXに変換し、推論するサンプル☆12Apr 11, 2024Updated 2 years ago
- Single-Image Depth Estimation Based on Fourier Domain Analysis☆16Feb 23, 2019Updated 7 years ago
- High performance Cross-platform Inference-engine, you could run Anakin on x86-cpu,arm, nv-gpu, amd-gpu,bitmain and cambricon devices.☆537Sep 23, 2022Updated 3 years ago
- Aidos Kuneen Full Node☆12Aug 18, 2022Updated 4 years ago
- Polyglot CUDA integration for the GraalVM☆18Apr 6, 2025Updated last year
- Sparse matrix-matrix multiplication on CPU+GPU systems.☆13Mar 17, 2014Updated 12 years ago
- ☆16Jan 16, 2023Updated 3 years ago
- ☆10Apr 23, 2021Updated 5 years ago
- High Performance Computing Conjugate Gradients: The original Mantevo miniapp☆20Jan 29, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Video classification using convGRU☆13Feb 15, 2018Updated 8 years ago
- Implementation of our CVPR2019 paper on Depth Completion: Dense Depth Posterior (DDP) from Single Image and Sparse Range☆17Mar 30, 2019Updated 7 years ago
- 用Paddle复现论文ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information(ACL2021)☆10Nov 15, 2021Updated 4 years ago
- ☆11Apr 10, 2015Updated 11 years ago
- ngAP's artifact for ASPLOS'24☆25Jul 29, 2025Updated last year
- An IR for efficiently simulating distributed ML computation.☆33Jan 13, 2024Updated 2 years ago
- A dynamic version of std::bitset☆17Aug 25, 2013Updated 13 years ago