Personal CUDA learning repo, built step by step from scratch.
☆104Jun 14, 2026Updated 2 months ago
Alternatives and similar repositories for Systematic-CUDA-Learning
Users that are interested in Systematic-CUDA-Learning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- EPOCH Input System Version 2☆10Jun 5, 2020Updated 6 years ago
- ☆56Jun 2, 2026Updated 2 months ago
- Cisco Live! BRKPRG-1798: Everybody Can Automate Now course material and reference☆11May 25, 2021Updated 5 years ago
- A lightweight, fully local-first PDF toolkit for macOS/iPhone/iPad☆36Jul 19, 2026Updated last month
- Source code with tasks from my "Write your own tiny programming system(s)!" course at Charles University. Follow the link below to watch…☆40Dec 6, 2025Updated 8 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Multi-class hallucination detection for LLM safety using contextual NLP classification. Supports experiment tracking with MLflow and mode…☆15Mar 10, 2026Updated 5 months ago
- A GPT-2 inference engine written from scratch in CUDA and C++. Implements custom CUDA kernels for tiled matrix multiplication, LayerNorm,…☆43May 17, 2026Updated 3 months ago
- Miscellaneous things and projects for my ZYBO and ZYNQ devices.☆11Aug 30, 2023Updated 3 years ago
- A PyTorch implementation of the GPT-OSS-20B architecture. All components are coded from scratch: RoPE with YaRN, RMSNorm, SwiGLU with cla…☆238Dec 2, 2025Updated 8 months ago
- ☆15Mar 11, 2026Updated 5 months ago
- Code repository accompanying O'Reilly MLOps with Databricks book☆35Aug 10, 2026Updated 2 weeks ago
- 🌐 A simple Python script designed to harness the power of Google Search in our quest for digital treasures.☆15Oct 7, 2024Updated last year
- a LLM inference engine to run on consumer hardware☆47Apr 15, 2026Updated 4 months ago
- ☆11Apr 22, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆28Updated this week
- ☆4,005Mar 11, 2026Updated 5 months ago
- A comprehensive Helm chart for monitoring GPU resources in Kubernetes clusters. This tool provides real-time visibility into GPU allocati…☆33Updated this week
- MegaDetector models served over FastAPI & visualized with Streamlit☆10May 9, 2023Updated 3 years ago
- ☆27Feb 16, 2024Updated 2 years ago
- Low-level C89 Closed/Open-addressed hashtable implementation.☆29Dec 28, 2024Updated last year
- ☆103Jan 29, 2026Updated 7 months ago
- Offline fuzzy geocoder — convert place names to coordinates with zero API limits☆56Feb 26, 2026Updated 6 months ago
- Minimum Viable Dataspace on AWS☆11Aug 5, 2026Updated 3 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆25Feb 13, 2026Updated 6 months ago
- A visual representation of Dijkstra's Algorithm using Libgdx.☆16Dec 24, 2021Updated 4 years ago
- A framework for self-improving prediction market trading agents.☆108Jul 15, 2026Updated last month
- Building a Computer From Scratch with verilog☆11Feb 6, 2026Updated 6 months ago
- ☆21May 11, 2026Updated 3 months ago
- The Westermo test system performance data set☆12Nov 24, 2023Updated 2 years ago
- Dockerfile for working with Modern Fortran☆27Mar 11, 2024Updated 2 years ago
- Enhancing the convergence speed by 2x and improving the training success of Physics-Informed Neural Networks (PINNs).☆13Oct 14, 2024Updated last year
- A useful tool for visualizing and analyzing where our models weaknesses are☆13Aug 15, 2019Updated 7 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Qwen3-0.6B megakernel: 527 tok/s decode on RTX 3090 (3.8x faster than PyTorch)☆135Feb 10, 2026Updated 6 months ago
- Meta-Reinforcement Learning with Self-Reflection☆34Mar 26, 2026Updated 5 months ago
- ☆10Jan 1, 2022Updated 4 years ago
- ☆15Apr 11, 2026Updated 4 months ago
- Structured close reading (or rather, close watching) transcripts of _almost _ every lesson in Jeremy Howard's Practical Deep Learning for…☆35Mar 9, 2026Updated 5 months ago
- An educational distributed training and inference library for neural nets using local computing☆77Jun 10, 2026Updated 2 months ago
- Fortran bindings to the C++ Standard Library.☆35Apr 7, 2025Updated last year