a whirlwind tour to deep learning and deep learning systems
☆81Jul 17, 2026Updated this week
Alternatives and similar repositories for ateenysitp
Users that are interested in ateenysitp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Hand-Rolled GPU communications library☆96Nov 25, 2025Updated 7 months ago
- ☆22Apr 22, 2024Updated 2 years ago
- NCCL communication API layer, and transport layer created from first principles.☆16Aug 20, 2025Updated 11 months ago
- Repo containing artifacts for Neurips 2025 tutorial- How to Build Agents to Generate Kernels for Faster LLMs (and Other Models!)☆15May 11, 2026Updated 2 months ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆231Jun 29, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- The best ChatGPT that $100 can buy.☆54Updated this week
- LLM RL envs done right, plus some training code☆17Feb 1, 2026Updated 5 months ago
- torch_remat fine-grained activation checkpointing API☆15Updated this week
- https://github.com/gpu-mode/reference-kernels☆26Jul 4, 2026Updated 2 weeks ago
- ☆19Dec 4, 2025Updated 7 months ago
- Code for the paper "Closing the Curious Case of Neural Text Degeneration"☆12Apr 9, 2025Updated last year
- MoE training for Me and You and maybe other people☆394Mar 15, 2026Updated 4 months ago
- Small scale distributed training of sequential deep learning models, built on Numpy and MPI.☆165Oct 19, 2023Updated 2 years ago
- Tutorials about tinygrad, an end-to-end deep learning stack☆97Jan 27, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆40Jun 26, 2026Updated 3 weeks ago
- Course on Flash-attention in Triton☆100Feb 9, 2026Updated 5 months ago
- minimal DL library in C: 24 NAIVE cuda/cpu ops, autodiff engine, python API (ops bindings/layers/models), tensor abstraction, strides, co…☆61Dec 17, 2025Updated 7 months ago
- ☆34Jun 28, 2026Updated 3 weeks ago
- A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do☆1,253Apr 27, 2026Updated 2 months ago
- An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.☆139Mar 21, 2026Updated 3 months ago
- ☆14Mar 29, 2026Updated 3 months ago
- 🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models☆215May 27, 2026Updated last month
- 100M tokens. Infinite compute. Lowest val loss wins.☆514Jul 3, 2026Updated 2 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Render MuJoCo scenes in bevy☆29Oct 7, 2025Updated 9 months ago
- This is a beginner-friendly tutorial on MLIR from the perspective of a user of MLIR, not a compiler engineer. This tutorial will introduc…☆140Mar 5, 2026Updated 4 months ago
- Write a fast kernel and see how you compare against the best humans and AI on gpumode.com☆103Updated this week
- Code for verifying deep neural feature ansatz☆22May 3, 2023Updated 3 years ago
- In fluid dynamics, an eddy is the swirling of a fluid and the reverse current created when the fluid is in a turbulent flow regime.☆18Updated this week
- Well documented examples of running distributed training jobs on Modal☆29Updated this week
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 10 months ago
- This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".☆139Sep 24, 2025Updated 9 months ago
- verl: Volcano Engine Reinforcement Learning for LLMs☆22Nov 6, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Training framework with a goal to explore the frontier of sample efficiency of small language models☆101Jan 25, 2026Updated 5 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆79Feb 18, 2026Updated 5 months ago
- ☆68Apr 8, 2026Updated 3 months ago
- Educational WIP☆73Feb 16, 2026Updated 5 months ago
- Code for "Approaching Deep Learning through the Spectral Dynamics of Weights"☆13Oct 30, 2024Updated last year
- ☆10Dec 17, 2019Updated 6 years ago
- Agentic RL Training at Scale☆1,696Updated this week