☆30Apr 18, 2024Updated 2 years ago
Alternatives and similar repositories for LibShalom
Users that are interested in LibShalom are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Apr 8, 2022Updated 4 years ago
- A direct convolution library targeting ARM multi-core CPUs.☆12Nov 27, 2024Updated last year
- Sparse kernels for GNNs based on TVM☆17Nov 18, 2020Updated 5 years ago
- Absinthe is an optimization framework to fuse and tile stencil codes in one shot☆14Jul 17, 2019Updated 7 years ago
- DietCode Code Release☆66Jul 21, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Apr 24, 2023Updated 3 years ago
- Stepwise optimizations of DGEMM on CPU, reaching performance faster than Intel MKL eventually, even under multithreading.☆165Feb 3, 2022Updated 4 years ago
- Spack package repository maintained by Student Cluster Competition Team @ Sun Yat-sen University.☆16Aug 20, 2025Updated last year
- A High performance and tiny TVM graph executor library written in C which can compile to WebAssembly and use CUDA/WebGPU as the accelerat…☆13Aug 3, 2023Updated 3 years ago
- Paper: inexact GMRES with fast multipole method and low-p relaxation☆11Aug 23, 2023Updated 3 years ago
- CSR5-based SpMV on CPUs, GPUs and Xeon Phi☆111Jun 10, 2024Updated 2 years ago
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 4 years ago
- HPC Challenge Benchmark☆71Sep 28, 2025Updated last year
- ☆10Jun 4, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆11Mar 2, 2024Updated 2 years ago
- ☆12May 3, 2020Updated 6 years ago
- This is the repo of "SEP-Graph: Finding Shortest Execution Paths for Graph Processing under a Hybrid Framework on GPU"☆16Dec 11, 2018Updated 7 years ago
- A Top-Down Profiler for GPU Applications☆24Feb 29, 2024Updated 2 years ago
- Python package to predict deep learning execution time☆13Jul 26, 2022Updated 4 years ago
- 慕课网 thinkphp5.0 微信小程序 零食商贩项目 小程序令牌测试工具☆12Dec 13, 2018Updated 7 years ago
- A comprehensive benchmarking framework for evaluating and optimizing CPU-centric agentic AI systems across multiple workloads, reproducin…☆56Oct 1, 2026Updated last week
- An attempt to replicate the paper "Multi-shot Pedestrian Re-identification via Sequential Decision Making (CVPR2018)"☆10Nov 16, 2019Updated 6 years ago
- study of cutlass☆22Nov 10, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Vector Math Library☆87Nov 7, 2025Updated 11 months ago
- NVFP4 Flash-Attention 4 on BlackWell☆68Sep 26, 2026Updated last week
- The repository maintains the source code for the article titled "Optimizing Attention by Exploiting Data Reuse on ARM Multi-core CPUs."☆17Dec 1, 2024Updated last year
- 中山大学2020年并行与分布式计算作业☆21Jul 28, 2020Updated 6 years ago
- CAKE Library for constant-bandwidth matrix multiplication on CPUs☆14Apr 6, 2024Updated 2 years ago
- Library for specialized dense and sparse matrix operations, and deep learning primitives.☆978Sep 28, 2026Updated last week
- Magicube is a high-performance library for quantized sparse matrix operations (SpMM and SDDMM) of deep learning on Tensor Cores.☆92Nov 23, 2022Updated 3 years ago
- 将MNN拆解的简易前向推理框架(for study!)☆24Feb 21, 2021Updated 5 years ago
- Document image binarization for Project 3A @Mines_Nancy☆30Aug 9, 2017Updated 9 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆23Aug 14, 2024Updated 2 years ago
- A source-to-source compiler for optimizing CUDA dynamic parallelism by aggregating launches☆15Jun 21, 2019Updated 7 years ago
- ☆2,043Jul 29, 2023Updated 3 years ago
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- GCN with MNE input and preprocessing☆10Sep 27, 2020Updated 6 years ago
- 我的毕业设计项目,一个能检测AI人脸合成图像的系统。☆11Apr 3, 2020Updated 6 years ago
- The code for paper: Neuralpower: Predict and deploy energy-efficient convolutional neural networks☆24Jul 10, 2019Updated 7 years ago