CUDA & Triton Learning Project: Flash Attention 实现探索
☆42Aug 14, 2025Updated last year
Alternatives and similar repositories for cuda-triton-learning
Users that are interested in cuda-triton-learning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Awesome code, projects, books, etc. related to CUDA☆43Aug 9, 2026Updated last month
- LeetGPU Solutions☆128Oct 9, 2025Updated 11 months ago
- Official implementation of the paper: "A deeper look at depth pruning of LLMs"☆15Jul 24, 2024Updated 2 years ago
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)☆31Jan 22, 2026Updated 7 months ago
- ☆10Mar 3, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- 基于CUDA的GPU加速通用遗传算法实现,实验平台为Nvidia Jetson Nano☆13Mar 23, 2023Updated 3 years ago
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 5 months ago
- ☆28Oct 11, 2022Updated 3 years ago
- Implementation code for ACL2024:Advancing Parameter Efficiency in Fine-tuning via Representation Editing☆15Apr 20, 2024Updated 2 years ago
- Solutions to leetgpu CUDA challenges on https://leetgpu.com/☆20May 25, 2025Updated last year
- 同济大学2019级数据库课程设计项目☆11Sep 11, 2021Updated 5 years ago
- Nano vLLM☆25Aug 11, 2025Updated last year
- An implementation of the penalty-based bilevel gradient descent (PBGD) algorithm and the iterative differentiation (ITD/RHG) methods.☆19Feb 13, 2023Updated 3 years ago
- Code for "Retaining Key Information under High Compression Rates: Query-Guided Compressor for LLMs" (ACL 2024)☆19Jun 12, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Blindspots in LLMs I've noticed while AI coding. Sonnet family emphasis.☆13Mar 20, 2025Updated last year
- A PyTorch implementation of computing mean average precision in parallel☆16Jul 7, 2022Updated 4 years ago
- Predictive Coding Network (PreCNet) for Next Frame Video Prediction.☆21Feb 13, 2023Updated 3 years ago
- 注释的nano_vllm仓库,并且完成了MiniCPM4的适配以及注册新模型的功能☆209Aug 11, 2025Updated last year
- ☆16Sep 12, 2023Updated 3 years ago
- A simple program scheduler for your code on different devices.☆12Mar 8, 2026Updated 6 months ago
- Inference with YOLOv5, OpenCV 4.5.4 DNN, C++, ROS and Python☆13Feb 12, 2023Updated 3 years ago
- ☆12Mar 7, 2024Updated 2 years ago
- Source code for our 2021 OpticsExpress paper "Robust super-resolution depth imaging via a multi-feature fusion deep network."☆22Nov 10, 2021Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆19Oct 6, 2020Updated 5 years ago
- common util library for C++☆12Updated this week
- Code of Robust Lottery Tickets for Pre-trained Language Models (ACL2022)☆20Jul 18, 2022Updated 4 years ago
- [EMNLP 2023]Context Compression for Auto-regressive Transformers with Sentinel Tokens☆25Nov 6, 2023Updated 2 years ago
- A self-learning tutorail for CUDA High Performance Programing.☆1,116Jan 14, 2026Updated 8 months ago
- Lottery Ticket Adaptation☆40Nov 20, 2024Updated last year
- ☆15Feb 3, 2021Updated 5 years ago
- Understanding the Linux 2.6.8.1 CPU Scheduler☆19Apr 24, 2015Updated 11 years ago
- ☆12Jan 19, 2020Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- RLVR training for LLM in CUDA/C++☆42Jul 25, 2026Updated last month
- 复旦研究生入学教育测试☆33Aug 28, 2025Updated last year
- 一步步通关GPU编程☆85Jun 4, 2026Updated 3 months ago
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆60Feb 6, 2026Updated 7 months ago
- Repository for the paper FFAVOD: Feature Fusion Architecture for Video Object Detection☆24Sep 21, 2021Updated 4 years ago
- Attack Active Directory Trusts with a single tool☆13Jan 15, 2025Updated last year
- 我会不断更新这个仓库中的文章☆13Jan 21, 2020Updated 6 years ago