a computing kernel implementation in ML inference framework aiming at theoretical limit
☆12Dec 18, 2019Updated 6 years ago
Alternatives and similar repositories for speedup-aarch64-cpu
Users that are interested in speedup-aarch64-cpu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A toy compiler for subset of c++ written in python☆16Jan 17, 2025Updated last year
- Kernel code for Samsung Galaxy S21 (Snapdragon 888)☆20Jul 4, 2021Updated 5 years ago
- A demo project for a computation graph implementation in C++.☆11Jul 2, 2019Updated 7 years ago
- An easy way to run, test, benchmark and tune OpenCL kernel files☆24Aug 25, 2023Updated 2 years ago
- C++ deepsort on tensorflow☆18Apr 4, 2020Updated 6 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A set of tools to work with cgroup tree and process classification/QoS according to it☆10Oct 1, 2019Updated 6 years ago
- LLVM passes with usage instructions☆18Apr 23, 2017Updated 9 years ago
- TensorFlow2.0 implementation FastFCN - https://arxiv.org/pdf/1903.11816v1.pdf☆11Aug 6, 2019Updated 6 years ago
- linux内核异步内存回收的另一个思路:基于冷热文件的冷热区域精准的回收冷文件页page(可做成内核ko)☆13Jun 14, 2024Updated 2 years ago
- 音视频分析工具☆12May 10, 2017Updated 9 years ago
- A Highlevel Python Wrapper for Vulkan's Compute API☆18Apr 13, 2026Updated 3 months ago
- A structure from motion implemention in C++ and accelerated using CUDA☆48Oct 12, 2019Updated 6 years ago
- Clone of https://code.google.com/p/google-coredumper/ with enhancements by Amadeus☆13Jul 2, 2024Updated 2 years ago
- The open-source project for "Mandheling: Mixed-Precision On-Device DNN Training with DSP Offloading"[MobiCom'2022]☆20Aug 4, 2022Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- C++ lock-free queue.☆14Jun 24, 2020Updated 6 years ago
- CUDA Template Functions☆20Dec 16, 2025Updated 7 months ago
- The repository targets the OpenCL gemm function performance optimization. It compares several libraries clBLAS, clBLAST, MIOpenGemm, Inte…☆17Mar 28, 2019Updated 7 years ago
- [IJCNN'19, IEEE JSTSP'19] Caffe code for our paper "Structured Pruning for Efficient ConvNets via Incremental Regularization"; [BMVC'18] …☆14Feb 14, 2020Updated 6 years ago
- Artifacts of EVT ASPLOS'24☆30Mar 6, 2024Updated 2 years ago
- Modified version of the YellowFin optimizer for TensorFlow to work with the Keras API [not actively maintained]☆16Jul 28, 2017Updated 9 years ago
- OpenDNN: An Open-source, cuDNN-like Deep Learning Primitive Library☆29Dec 9, 2019Updated 6 years ago
- Documentation for the entire CGRAFlow☆19Sep 17, 2021Updated 4 years ago
- OSDT2019相关资料☆16Nov 17, 2019Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- dnotify,inotify, and fanotify example code from http://www.lanedo.com/filesystem-monitoring-linux-kernel/☆14Apr 28, 2017Updated 9 years ago
- dump zmq messages on a socket☆15Aug 27, 2023Updated 2 years ago
- A simple cycle-accurate DaDianNao simulator☆13Mar 27, 2019Updated 7 years ago
- 一个尝试固液耦合的沙盒玩具☆11Feb 17, 2025Updated last year
- ☆14Dec 8, 2022Updated 3 years ago
- convert pytorch trained yolo model to ncnn for Flexible deployment☆10Aug 30, 2018Updated 7 years ago
- Project for testing remote debugging of C++ code with gdb and gdbserver in VS Code☆20Jun 6, 2018Updated 8 years ago
- C/C++ header dependency list generator. Output can be used to create a dependency graph.☆15Apr 13, 2021Updated 5 years ago
- ☆16Aug 11, 2016Updated 9 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- NVM user-space Primitives API library repository☆18Mar 12, 2014Updated 12 years ago
- Basic algorithms for vslam.☆54Nov 20, 2020Updated 5 years ago
- This is a pintool that can analyze target dynamically and output code blocks and "key frames".☆14Mar 26, 2015Updated 11 years ago
- SAF: Streaming Analytics Framework☆31Mar 6, 2019Updated 7 years ago
- log, 仅包含头文件,追踪崩溃和数据的日志库☆16Dec 25, 2018Updated 7 years ago
- Linux热补丁实践☆18Jun 11, 2019Updated 7 years ago
- ☆19Jul 17, 2026Updated last week