☆40Jun 25, 2026Updated last month
Alternatives and similar repositories for hpc2torch
Users that are interested in hpc2torch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- easy cuda code☆101Dec 24, 2024Updated last year
- some hpc project for learning☆28Aug 28, 2024Updated last year
- 训练营讲义☆21Jan 21, 2025Updated last year
- 笔记☆53Jul 3, 2026Updated 3 weeks ago
- ☆80Jan 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Flash Attention in ~100 lines of CUDA (forward pass only)☆12Jun 10, 2024Updated 2 years ago
- ☆17Mar 29, 2026Updated 3 months ago
- Some funny cute/cuteDSL code snippets☆33Mar 2, 2026Updated 4 months ago
- 南京大学2022春季PA实验☆13Aug 27, 2023Updated 2 years ago
- "aura" my super-scalar O3 cpu core☆26May 25, 2024Updated 2 years ago
- ACPO: Python ML training/inference framework for LLVM optimizations☆13Jul 2, 2025Updated last year
- Large-scale Auto-Distributed Training/Inference Unified Framework | Memory-Compute-Control Decoupled Architecture | Multi-language SDK & …☆54Jul 17, 2026Updated last week
- ☆13Aug 9, 2022Updated 3 years ago
- simplify >2GB large onnx model☆72Nov 30, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implementation of BitonicSorting algorithm on FPGA through SDAccel using Opencl as source code☆17Nov 21, 2016Updated 9 years ago
- Public repostory for the DAC 2021 paper "Scaling up HBM Efficiency of Top-K SpMV forApproximate Embedding Similarity on FPGAs"☆16Aug 29, 2021Updated 4 years ago
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated 11 months ago
- practical C programming skill☆53Jan 16, 2025Updated last year
- GPU-accelerated linear solvers based on the conjugate gradient (CG) method, supporting NVIDIA and AMD GPUs with GPU-aware MPI, NCCL, RCCL…☆16Mar 14, 2026Updated 4 months ago
- ☆149Aug 18, 2025Updated 11 months ago
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆23Apr 10, 2026Updated 3 months ago
- ☆28Aug 9, 2025Updated 11 months ago
- EasyNN是一个面向教学而开发的神经网络推理框架,旨在让大家0基础也能自主完成推理框架编写!☆40Aug 26, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- fake CUTLASS to get peformance☆26Apr 28, 2026Updated 2 months ago
- 中文版 LLM101n 课程☆71Aug 11, 2024Updated last year
- ☆52Mar 4, 2026Updated 4 months ago
- ☆44Jan 8, 2025Updated last year
- ☆11Nov 2, 2017Updated 8 years ago
- InfiniTensor 大模型与人工智能系统训练营 CUDA 方向作业与项目系统☆51Feb 24, 2026Updated 5 months ago
- LLVM OpenCL C compiler suite for ventus GPGPU☆64Jul 17, 2026Updated last week
- 基4booth乘法器设计与验证☆15Apr 28, 2024Updated 2 years ago
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.☆21Jan 24, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- High Performance LLM Inference Operator Library☆1,065Updated this week
- Code for the "Evolving Reservoirs for Meta Reinforcement Learning" paper☆12Apr 22, 2024Updated 2 years ago
- ☆32Jul 2, 2025Updated last year
- 基于c++ muduo网络库的集群聊天服务器,使用nginx实现负载均衡,使用reids消息队列实现跨服务器通信☆13Feb 23, 2024Updated 2 years ago
- SIGIR 2023 "Graph Collaborative Signals Denoising and Augmentation for Recommendation"☆20May 31, 2023Updated 3 years ago
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 3 months ago
- CUDA C simple application for Nvidia's GPU☆11Jun 7, 2022Updated 4 years ago