Portable and Flexible DGEMM Library for GPUs (OpenCL, CUDA, CAL) with special support for HPL
☆16Apr 5, 2018Updated 8 years ago
Alternatives and similar repositories for caldgemm
Users that are interested in caldgemm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High Performance Linpack for GPUs (Using OpenCL, CUDA, CAL)☆93Oct 22, 2015Updated 10 years ago
- GPU implementation of classical molecular dynamics proxy application.☆31Jan 30, 2017Updated 9 years ago
- Gamebaby Rock Sun's D3D12 DirectX Ray Tracing C-Style Sample for beginner☆18Feb 5, 2023Updated 3 years ago
- An HPL-AI implementation for Fugaku☆24Jun 29, 2021Updated 5 years ago
- Argonne Leadership Computing Facility OpenCL tutorial☆10Aug 22, 2025Updated 11 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Made with LaTex. NENU's recommendation letter template.☆12May 26, 2024Updated 2 years ago
- OpenCL porting of the GROMACS molecular simulation toolkit☆27Sep 5, 2015Updated 10 years ago
- ☆11Aug 8, 2021Updated 4 years ago
- A neutral particle transport mini-app to study performance of sweeps on unstructured, 3D tetrahedral meshes.☆19Sep 20, 2022Updated 3 years ago
- Rapid HPC Orchestration in the Cloud☆28Oct 3, 2023Updated 2 years ago
- Open source of an IBM Optimized version of the HPCG benchmark.☆17Sep 17, 2025Updated 10 months ago
- the proxy that breaks all the rules and doesn't ever care.☆16Jun 18, 2012Updated 14 years ago
- EPCC I/O benchmarking applications☆12Dec 15, 2021Updated 4 years ago
- GPU Optimization and Memory Abstraction Framework☆33Oct 31, 2019Updated 6 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A python script that reads in a fortran 77 (.f or .F) fixed form file and converts it to a free form Fortran 90 file (.f90 or .F90).☆25Apr 12, 2016Updated 10 years ago
- Tools to run and parse MKL verbose mode☆18Jun 28, 2022Updated 4 years ago
- JIT Compilation for Multiple Architectures: C++, OpenMP, CUDA, HIP, OpenCL, Metal☆13Jun 10, 2026Updated last month
- Phonak - Hearing Aid App: a solution enabling patients to receive hearing aid adjustments remotely via a Bluetooth, Internet-enabled smar…☆12Jul 27, 2014Updated 11 years ago
- EPCC OpenACC Benchmarks☆19Sep 23, 2013Updated 12 years ago
- Source of BLAS via BLIS☆13May 24, 2026Updated 2 months ago
- Modern Fortran wrappers around MPI routines☆36Dec 17, 2025Updated 7 months ago
- Multigrid Methods - An Overview, A lecture series at Imperial College☆26Jan 6, 2025Updated last year
- C and NVidia CUDA code for multi-view deconvolution☆12Jul 25, 2016Updated 9 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ROCm Command Line Profiler - Updated moved to https://github.com/GPUOpen-Tools/RCP☆10Aug 24, 2017Updated 8 years ago
- C++1y coroutine library.☆17Mar 26, 2018Updated 8 years ago
- DEPRECATED: GlovePIE scripts for playing Descent with gamepads, such as the PS3 and XBOX 360 controllers.☆19Jun 22, 2022Updated 4 years ago
- GPU Debugging SDK for ROCm☆10Mar 21, 2019Updated 7 years ago
- QUICK, a GPU-enabled ab intio quantum chemistry software. Now move to the main branch: https://github.com/merzlab/QUICK☆11Jan 19, 2015Updated 11 years ago
- Update dynamic DNS records from netlink☆15Mar 24, 2026Updated 3 months ago
- Compute applications.☆25Dec 12, 2019Updated 6 years ago
- ☆14Aug 4, 2022Updated 3 years ago
- Introduction to OpenACC☆30Jan 25, 2021Updated 5 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A tool allowing students of Coursera's Heterogeneous Parallel Programming to work on homework using a machine without a CUDA GPU.☆11Mar 11, 2015Updated 11 years ago
- A Monte Carlo Neutron Transport Mini-App☆15Apr 15, 2019Updated 7 years ago
- C library containing high resolution timer implementation for several platforms.☆10Oct 20, 2020Updated 5 years ago
- Neural Network Extension for Ruby☆21Dec 1, 2011Updated 14 years ago
- Code examples for the CUDA workshop☆36Sep 19, 2022Updated 3 years ago
- How we hooked &yet's SimpleWebRTC into XirSys's STUN and TURN servers☆19Jul 16, 2016Updated 10 years ago
- Workflow management system for the automated and distributed analysis of large-scale experimental data.☆13Oct 3, 2024Updated last year