pure c/cpp cnn implementation, with CUDA accelerated.
☆21Apr 30, 2021Updated 5 years ago
Alternatives and similar repositories for SimpleCNN_Release
Users that are interested in SimpleCNN_Release are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- C++ implement a simple CNN framework to train mnist data. Done!☆10Mar 29, 2022Updated 4 years ago
- 方便扩展的Cuda算子理解和优化框架,仅用在学习使用☆18Jun 13, 2024Updated 2 years ago
- AST interpreter with clang 5.0.0 and llvm 5.0.0☆14Dec 7, 2019Updated 6 years ago
- Official implementation of Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores.☆18Nov 13, 2025Updated 9 months ago
- [AAAI 2026] This is the official implementation of the paper "ExtendAttack: Attacking Servers of LRMs via Extending Reasoning".☆26Mar 18, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Project dedicated to making GPU Partitioning on Windows easier!☆15Jan 10, 2022Updated 4 years ago
- Implementation of the D-Stream clustering algorithm for use in MOA. An earlier version is included as part of the MOA 17.06 release.☆14Jan 29, 2019Updated 7 years ago
- CNN accelerated by cuda. Test on mnist and finilly get 99.76%☆188Oct 15, 2017Updated 8 years ago
- 国科大编译作业:基于Clang的C语言解释执行器☆43Dec 12, 2021Updated 4 years ago
- ☆11Sep 12, 2023Updated 2 years ago
- A D-Stream clustering algorithm implementation in Python☆14Mar 25, 2016Updated 10 years ago
- ☆11Oct 9, 2019Updated 6 years ago
- 🏆 The 1st Place Solution for AICity2022 Challenge Track2: Natural Language-Based Vehicle Retrieval.☆12Jul 25, 2022Updated 4 years ago
- ☆15Dec 6, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Build CUDA Neural Network From Scratch☆22Aug 28, 2024Updated last year
- ☆20Sep 28, 2024Updated last year
- OpenCL implementation of BM3D image denoising algorithm☆11Oct 28, 2019Updated 6 years ago
- This repo is "NTHU Parallel Programing" course project.☆10Dec 5, 2017Updated 8 years ago
- Opara is a lightweight and resource-aware DNN Operator parallel scheduling framework to accelerate the execution of DNN inference on GPUs…☆23Dec 19, 2024Updated last year
- The 1st Place Submission to AICity Track5 - Natural Language-based Vehicle Retrieval.☆16Aug 14, 2021Updated 5 years ago
- Compress BiSeNet with Structure Knowledge Distillation for Real-time image segmentation on wali-TX2☆11Jul 29, 2020Updated 6 years ago
- FlashSparse significantly reduces the computation redundancy for unstructured sparsity (for SpMM and SDDMM) on Tensor Cores through a Swa…☆39Oct 5, 2025Updated 10 months ago
- NTHU CS6135 VLSI實體設計自動化☆11Mar 12, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Fastest CUDA SIFT or other 128-float vector matcher for computer vision☆28Mar 23, 2021Updated 5 years ago
- 适配目前最新CUDA环境的SIFTGPU代码☆10May 6, 2020Updated 6 years ago
- A fast, small, efficient pthreads based threadpool in c☆16Mar 2, 2021Updated 5 years ago
- a simple API to use CUPTI☆10Aug 19, 2025Updated 11 months ago
- 一步步实现c++中的智能指针☆10Jun 6, 2021Updated 5 years ago
- Minimal RISC-V Chisel design strictly reflecting the ISA document for verification.☆21Apr 14, 2026Updated 4 months ago
- ☆11Sep 13, 2020Updated 5 years ago
- Triton to TVM transpiler.☆24Oct 14, 2024Updated last year
- 'Efficient High-Frequency Texture Recovery Diffusion Model for Remote Sensing Image Super-Resolution'☆17Dec 11, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 计算机网络微课堂笔记☆16Jul 1, 2023Updated 3 years ago
- Fastest CUDA RGB to grayscale: 5-30x faster than OpenCV. For image processing/computer vision.☆16Mar 23, 2021Updated 5 years ago
- [ACM MM 2024] Mutual-Guided Dynamic Network for Image Fusion☆21Aug 24, 2023Updated 2 years ago
- ☆23Mar 12, 2026Updated 5 months ago
- An ATPG tool using PODEM algorithm in C++ that generates a test to detect any given list of Single-Stuck-at Faults☆11Oct 29, 2017Updated 8 years ago
- Implementation of a simple CNN using CUDA☆68May 2, 2017Updated 9 years ago
- High Performance Grouped GEMM in PyTorch☆30May 10, 2022Updated 4 years ago