AI Infrastructure Performance Engineer Learning Track - GPU optimization, inference optimization, and cost reduction
☆76Jul 8, 2026Updated 3 months ago
Alternatives and similar repositories for ai-infra-performance-learning
Users that are interested in ai-infra-performance-learning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the WSDM 2021 paper "FluxEV: A Fast and Effective Unsupervised Framework for Time-Series Anomaly Detection".☆22Mar 19, 2024Updated 2 years ago
- AI Infrastructure Engineer Learning Track - Production ML infrastructure curriculum (2-4 years experience)☆1,769Jun 26, 2026Updated 3 months ago
- ☆14Aug 7, 2024Updated 2 years ago
- Project Repo for GSoC 2019 at CERN☆11Jan 4, 2023Updated 3 years ago
- Yandex Style Guide☆11Sep 10, 2015Updated 11 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- learn TensorRT from scratch🥰☆18Sep 29, 2024Updated 2 years ago
- LLM KV Cache compression - K+V dual compression, 73-99% VRAM savings, zero accuracy loss☆57Mar 30, 2026Updated 6 months ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆45Jul 22, 2026Updated 2 months ago
- GEMM☆10Aug 26, 2023Updated 3 years ago
- A Chinese-focused PyTorch framework for exploring Attention Residuals in Qwen3-style causal LMs, with baseline, Block AttnRes, Full AttnR…☆21May 3, 2026Updated 5 months ago
- This comprehensive learning repository is designed to transform software engineers into expert AI kernel developers, focusing on the cutt…☆78Mar 25, 2026Updated 6 months ago
- General, reusable APIs and code used by the various Submariner components.☆28Updated this week
- ☆11May 16, 2026Updated 4 months ago
- (Academia Sinica / Computer Vision / Deep Learning) Object Detection, Person Reid, Face Reid☆13Nov 21, 2022Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Pipeline for text-classification (binary and multiclass problem)☆15Aug 5, 2020Updated 6 years ago
- Dash app with animated scatter map on mapbox/plotly (COVID data)☆13Apr 17, 2024Updated 2 years ago
- 🎓Automatically Update circult-eda-mlsys-tinyml Papers Daily using Github Actions (Update Every 8th hours)☆10Updated this week
- My tests and experiments with some popular dl frameworks.☆17Sep 11, 2025Updated last year
- Google STEP Internship dev course - TSP Challenges☆11Jul 7, 2017Updated 9 years ago
- ☆14Nov 3, 2025Updated 11 months ago
- ☆15Sep 8, 2022Updated 4 years ago
- 52 weeks, 52 projects☆11Apr 28, 2024Updated 2 years ago
- Multi-heap-sort for many small arrays, quicksort with 3 pivots for one big array, CUDA acceleration, CUDA memory compression.☆13Sep 29, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Become GPU kernel engineer step by step.☆72May 7, 2026Updated 5 months ago
- Design and Analysis of Algorithms - 1 (from Stanford University)☆19Nov 6, 2018Updated 7 years ago
- The AI-Native SDLC Playbook — Anthropic Claude Academy course as EPUB ebook☆93Sep 3, 2026Updated last month
- An automatically annotated sentiment analysis dataset of product reviews in Russian.☆17Oct 25, 2020Updated 5 years ago
- WWCode AWS Certification Study Group Resources☆14Dec 5, 2020Updated 5 years ago
- Mathematical consequences of orthogonal weights initialization and regularization in deep learning. Experiments with gain-adjusted orthog…☆17Sep 21, 2019Updated 7 years ago
- My thesis project for person re-identification☆14Dec 8, 2022Updated 3 years ago
- DukeMTMC-reID_baseline (Matlab)☆18Aug 8, 2017Updated 9 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Fast GPU based tensor core reductions☆12Jan 13, 2023Updated 3 years ago
- portFFT is a library implementing Fast Fourier Transforms using SYCL☆19Mar 1, 2025Updated last year
- Make FP4 on 5090 Great Again☆24Jul 20, 2026Updated 2 months ago
- ☆13Aug 31, 2023Updated 3 years ago
- a reactor network library☆16Aug 21, 2025Updated last year
- Skeleton-based method for activity recognition problem☆13Dec 20, 2022Updated 3 years ago
- Batched square compact-Householder QR factorization.☆15Jul 2, 2026Updated 3 months ago