tutorial to optimize GEMM performance on android
☆51Feb 17, 2016Updated 10 years ago
Alternatives and similar repositories for gemm-android
Users that are interested in gemm-android are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- profiling gemm on android☆10Apr 1, 2016Updated 10 years ago
- Low-precision matrix multiplication☆1,845Jan 29, 2024Updated 2 years ago
- ICME 2016 "Learning Deep Representation from Coarse to Fine for Face Alignment"☆30Oct 29, 2018Updated 7 years ago
- ☆11Sep 10, 2025Updated 11 months ago
- A faster re-implementation of the FAST-9 algorithm (C++, with C bindings available)☆14Feb 1, 2017Updated 9 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Kernel Fusion and Runtime Compilation Based on NNVM☆72Nov 21, 2016Updated 9 years ago
- SqueezeNet: AlexNet-level accuracy with 50x fewer parameters☆21Mar 2, 2016Updated 10 years ago
- The benchmark of ncnn that is a high-performance neural network inference framework optimized for the mobile platform☆72Mar 8, 2019Updated 7 years ago
- A program that times various techniques for performing a moving median filter (sometimes called rolling median, or streaming median)☆11Feb 13, 2016Updated 10 years ago
- Porting caffe to android platform☆10Jul 16, 2016Updated 10 years ago
- ☆20Dec 15, 2023Updated 2 years ago
- Train Neuronal networks to automate your home☆21Mar 1, 2023Updated 3 years ago
- This is an read-only mirror of the gem5 simulator. The upstream repository is stored in https://gem5.googlesource.com, code reviews shoul…☆19Aug 21, 2021Updated 4 years ago
- Greentea LibDNN - a universal convolution implementation supporting CUDA and OpenCL☆137Apr 20, 2017Updated 9 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A simple baseline model set using MXNet for Kaggle StateFarm driver position identification☆27Jul 1, 2016Updated 10 years ago
- CLTune: An automatic OpenCL & CUDA kernel tuner☆186Dec 12, 2022Updated 3 years ago
- Deep CNN on Android☆30Feb 26, 2017Updated 9 years ago
- Acceleration package for neural networks on multi-core CPUs☆1,710Jun 11, 2024Updated 2 years ago
- Open Source Library for GPU-Accelerated Execution of Trained Deep Convolutional Neural Networks on Android☆540Apr 12, 2017Updated 9 years ago
- Proof-of-Concept CNN in Halide☆22Aug 4, 2016Updated 10 years ago
- Optimizing Mobile Deep Learning on ARM GPU with TVM☆184Oct 15, 2018Updated 7 years ago
- NNVM for ROCm Examples☆19Nov 22, 2017Updated 8 years ago
- a software library containing BLAS functions written in OpenCL☆866Aug 2, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Amalgamation and go binding☆63Nov 11, 2015Updated 10 years ago
- Personal collection of references for high performance mixed precision training.☆41Oct 21, 2019Updated 6 years ago
- Neural Style Transfer with Caffe2 on your Android phone☆82Mar 28, 2019Updated 7 years ago
- The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologi…☆3,180Updated this week
- Portable 128-bit SIMD intrinsics☆61Jul 4, 2023Updated 3 years ago
- C++ training and testing code for an SVM using Vlfeat fisher vectors together with possible other features.☆14Jun 29, 2016Updated 10 years ago
- ☆17Aug 22, 2021Updated 4 years ago
- Automatically exported from code.google.com/p/math-neon☆40Apr 20, 2015Updated 11 years ago
- ☆402Mar 15, 2019Updated 7 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆10May 4, 2023Updated 3 years ago
- CUDA Extension Wrangler☆26Aug 21, 2019Updated 6 years ago
- Winograd minimal convolution algorithm generator for convolutional neural networks.☆629Feb 9, 2026Updated 6 months ago
- Communication-Minimizing 2D Convolution in GPU Registers☆30Sep 21, 2013Updated 12 years ago
- A set of benchmarks to compare the main 2D marker detection and tracking libraries☆11Oct 19, 2017Updated 8 years ago
- Porting caffe to android platform☆506Dec 11, 2018Updated 7 years ago
- BLISlab: A Sandbox for Optimizing GEMM☆572Jun 17, 2021Updated 5 years ago