🎉CUDA 笔记 / 高频面试题汇总 / C++笔记,个人笔记,更新随缘: sgemm、sgemv、warp reduce、block reduce、dot product、elementwise、softmax、layernorm、rmsnorm、hist etc.
☆52Feb 23, 2024Updated 2 years ago
Alternatives and similar repositories for CUDA-Learn-Note
Users that are interested in CUDA-Learn-Note are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for paper: Latent-space Dynamics for Reduced Deformable Simulation☆39May 29, 2019Updated 7 years ago
- A 3D fluid simulation on the GPU using C++ and Vulkan.☆13Jun 12, 2022Updated 4 years ago
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 4 years ago
- ☆14Apr 16, 2024Updated 2 years ago
- Adaptive Topology Reconstruction for Robust Graph Representation Learning [Efficient ML Model]☆10Feb 11, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆23Aug 14, 2024Updated 2 years ago
- An implementation of "Air Meshes for Robust Collision Handling", SIGGRAPH (2015)☆16Aug 30, 2017Updated 9 years ago
- 从dpdk里面扒下来的ret_ring☆18Oct 27, 2014Updated 11 years ago
- NeurIPS 2020 Spotlight Paper☆13Dec 20, 2021Updated 4 years ago
- ☆11Jun 13, 2022Updated 4 years ago
- 🎉CUDA 笔记 / 高频面试题汇总 / C++笔记,个人笔记,更新随缘: sgemm、sgemv、warp reduce、block reduce、dot product、elementwise、softmax、layernorm、rmsnorm、hist etc.☆51Jan 25, 2024Updated 2 years ago
- Implementation of Speculative Sampling as described in "Accelerating Large Language Model Decoding with Speculative Sampling" by Deepmind☆112Feb 29, 2024Updated 2 years ago
- ARP4G是一个go语言实现的简化应用开发的框架。使开发者专注于产品业务本身。☆21Nov 29, 2022Updated 3 years ago
- A large-scale training and benchmarking framework for rPPG.☆10Nov 26, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆23May 10, 2023Updated 3 years ago
- Implementation of our paper: Komaritzan and Botsch, Fast Projective Skinning, ACM MIG 2019.☆59Jan 27, 2024Updated 2 years ago
- ☆17Dec 29, 2025Updated 8 months ago
- Linear Attention for Efficient Bidirectional Sequence Modeling☆18May 13, 2025Updated last year
- ☆13Apr 17, 2024Updated 2 years ago
- Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.☆11,900Updated this week
- ☆13May 14, 2024Updated 2 years ago
- Multi-GPU Framework for Voxel Grid Computations☆69Jul 7, 2026Updated 2 months ago
- A C++ port of karpathy/micrograd, a tiny scalar-valued autograd engine and a neural net library☆13Nov 24, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- taichi hackathon repo.☆18Dec 15, 2022Updated 3 years ago
- A Winograd Minimal Filter Implementation in CUDA☆31Aug 25, 2021Updated 5 years ago
- Subspace Graph Physics☆13Jun 14, 2024Updated 2 years ago
- some physics implemented on Taichi-AOT & Unity☆17Dec 4, 2022Updated 3 years ago
- CoRdE model implementation: simulating ropes, chains, and other elastic strings☆11May 8, 2020Updated 6 years ago
- SuperTerrain+: A real-time procedural 3D infinite terrain engine with geographical features and photorealistic rendering.☆18Apr 6, 2023Updated 3 years ago
- molecular dynamics (MD) simulation of 10^13 atoms.☆12Nov 22, 2024Updated last year
- Official implementation of "TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization" (Findings of ACL …☆21Jul 25, 2025Updated last year
- ☆11Jan 3, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Render a controller overlay during gameplay.☆17Nov 16, 2023Updated 2 years ago
- CUDA project for uni subject☆26Oct 26, 2020Updated 5 years ago
- Parallel Prefix Sum (Scan) with CUDA☆30Jun 22, 2024Updated 2 years ago
- Modified g2o with GPU support for general matrix calculations.☆16Jun 28, 2025Updated last year
- ☆15Aug 29, 2021Updated 5 years ago
- ☆13May 18, 2022Updated 4 years ago
- The OFELI Library: An Object Oriented Finite Element Library☆18Jan 8, 2026Updated 8 months ago