☆52Mar 4, 2026Updated 4 months ago
Alternatives and similar repositories for triton_course
Users that are interested in triton_course are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A light llama-like llm inference framework based on the triton kernel.☆188Jan 5, 2026Updated 6 months ago
- llm theoretical performance analysis tools and support params, flops, memory and latency analysis.☆119Jul 11, 2025Updated last year
- 校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。☆555Oct 28, 2025Updated 9 months ago
- This project is primarily used to deploy large language models and multimodal large models on Orin.🚀🚀🚀☆18Jun 23, 2026Updated last month
- A CUDA tutorial to make people learn CUDA program from 0☆279Jul 9, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Triton Migration Guide for DeepStreamSDK.☆15Dec 19, 2023Updated 2 years ago
- CUDA SGEMM optimization note☆15Oct 31, 2023Updated 2 years ago
- 校招、秋招、春招、实习好项目!带你从零实现一个高性能的深度学习推理库,支持大模型 llama2 、Unet、Yolov5、Resnet等模型的推理。Implement a high-performance deep learning inference library st…☆3,476Jun 22, 2025Updated last year
- CUDA 13.1 Tutorial Series for RTX 5090 (Blackwell) - Chinese teaching materials☆29Jan 18, 2026Updated 6 months ago
- Implement Flash Attention using Cute.☆111Dec 17, 2024Updated last year
- ☆234Updated this week
- some hpc project for learning☆28Aug 28, 2024Updated last year
- A skill for automatically optimizing CUDA code.☆42Mar 26, 2026Updated 4 months ago
- ☆65Apr 26, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆13Jul 28, 2024Updated 2 years ago
- TensorRT encapsulation, learn, rewrite, practice.☆31Oct 19, 2022Updated 3 years ago
- minimal Vision Language Action framework for robot control systems☆17Sep 15, 2025Updated 10 months ago
- YOLO Series☆14Oct 20, 2023Updated 2 years ago
- my cs notes☆72Oct 14, 2024Updated last year
- ☆17Apr 23, 2026Updated 3 months ago
- 使用 CUDA C++ 实现的 llama 模型推理框架☆65Nov 8, 2024Updated last year
- ☆13Oct 17, 2024Updated last year
- 本仓库在OpenVINO推理框架下部署Nanodet检测算法,并重写预处理和后处理部分,具有超高性能!让你在Intel CPU平台上的检测速度起飞! 并基于NNCF和PPQ工具将模型量化(PTQ)至int8精度,推理速度更快!☆16Jun 14, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- This is a series of GPU optimization topics. Here we will introduce how to optimize the CUDA kernel in detail. I will introduce several…☆1,335Jul 29, 2023Updated 3 years ago
- The official implement of CTRNet++.☆15Dec 30, 2024Updated last year
- ☆12May 19, 2022Updated 4 years ago
- ゼロから作るDeep Learning ❸ をC++で実装する。自習用リポジトリ。☆18Aug 12, 2020Updated 5 years ago
- This repository contains my coursework and projects completed during the GPU Programming Specialization offered by Johns Hopkins Universi…☆11Jun 13, 2023Updated 3 years ago
- ☆40Jun 25, 2026Updated last month
- ☆150Jan 9, 2025Updated last year
- ☆19Nov 10, 2024Updated last year
- CPU Memory Compiler and Parallel programing☆26Nov 18, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Accelerating SAHI-based inference on YOLO models using TensorRT.☆103Jan 6, 2026Updated 6 months ago
- LLM Inference via Triton (Flexible & Modular): Focused on Kernel Optimization using CUBIN binaries, Starting from gpt-oss Model☆119Apr 28, 2026Updated 3 months ago
- https://github.com/shouxieai/hard_decode_trt windows编译版本☆13Sep 8, 2022Updated 3 years ago
- ☆23Aug 20, 2025Updated 11 months ago
- ☆15Oct 9, 2022Updated 3 years ago
- CS149 xmake version☆46Nov 30, 2023Updated 2 years ago
- Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruct…☆558Sep 8, 2024Updated last year