GPT2 in handwritten PTX
☆16Jun 29, 2025Updated last year
Alternatives and similar repositories for llm-ptx
Users that are interested in llm-ptx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- a platform for monitoring the chip situation☆16Jul 19, 2025Updated last year
- ☆47Jul 27, 2026Updated 2 weeks ago
- Parallel Associative Scan for Language Models☆18Jan 8, 2024Updated 2 years ago
- A library for constrained RLHF.☆13Feb 19, 2024Updated 2 years ago
- High-Performance FP32 GEMM on CUDA devices☆126Jan 21, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Starlight: A Kernel Optimizer for GPU Processing☆16Jan 10, 2024Updated 2 years ago
- ☆13Updated this week
- A Binary Translation Framework for CUDA☆24Aug 7, 2014Updated 12 years ago
- FAST Randomized SVD on a GPU with CUDA 🏎️☆16May 21, 2019Updated 7 years ago
- ☆21Oct 19, 2022Updated 3 years ago
- ☆15Feb 23, 2025Updated last year
- Collection of useful kernels☆20Oct 3, 2025Updated 10 months ago
- This is the official repository for "CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data An…☆24Oct 26, 2023Updated 2 years ago
- Programming an RTS by Carl Granberg.☆15Apr 23, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A version of nuttx used by smoothie-v2☆10Apr 1, 2018Updated 8 years ago
- ☆30Apr 7, 2025Updated last year
- An app to compute the coefficients of a function development in a spherical harmonics convergent series.☆19Sep 14, 2022Updated 3 years ago
- [EMNLP 24] Source code for paper 'AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tu…☆13Dec 15, 2024Updated last year
- ☆19Nov 6, 2024Updated last year
- QuintNet is a research-oriented PyTorch framework designed to explore and implement multi-dimensional parallelism strategies for distribu…☆19Feb 3, 2026Updated 6 months ago
- a simple Flash Attention v2 implementation with ROCM (RDNA3 GPU, roc wmma), mainly used for stable diffusion(ComfyUI) in Windows ZLUDA en…☆54Aug 25, 2024Updated last year
- A quick way to render multiple OpenSCAD files using python3's thread support.☆14Jun 26, 2025Updated last year
- Breakout board for ICM-20948 with i2c☆17Feb 16, 2019Updated 7 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Distributed MoE in a Single Kernel [NeurIPS '25]☆281May 5, 2026Updated 3 months ago
- Experimentation in decoding serial UART for the ES200G battery—only the battery, not the scooter's communications to it.☆12Oct 30, 2022Updated 3 years ago
- Course Project for COMP4471 on RWKV☆17Feb 11, 2024Updated 2 years ago
- A system to manage online orders across Amazon, Ebay, Walmart, Reverb, and Big Commerce stores☆14Mar 13, 2023Updated 3 years ago
- ☆25Oct 10, 2025Updated 10 months ago
- ☆82Dec 27, 2024Updated last year
- Concise neural network with C++ and CUDA☆16Apr 16, 2021Updated 5 years ago
- Public Source code Release of Theori's AIxCC ASC Submission☆15Aug 5, 2025Updated last year
- DO NOT USE THIS PROJECT - WILL BE SUPERCEDED☆17Apr 20, 2018Updated 8 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- NCCL communication API layer, and transport layer created from first principles.☆16Aug 20, 2025Updated 11 months ago
- ☆37Jun 10, 2025Updated last year
- Write a fast kernel and see how you compare against the best humans and AI on gpumode.com☆109Jul 29, 2026Updated last week
- ☆18Apr 8, 2025Updated last year
- Using Amazon Alexa Echo device to control ESP32 device☆18Jun 12, 2017Updated 9 years ago
- The simplest but fast implementation of matrix multiplication in CUDA.☆39Jul 26, 2024Updated 2 years ago
- Kendryte k210 freertos programming guide☆19Oct 18, 2019Updated 6 years ago