a minimal paged attention implementation
☆20Jan 30, 2026Updated 6 months ago
Alternatives and similar repositories for nano-paged-attention
Users that are interested in nano-paged-attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- writing really fast kernels☆19Jul 15, 2026Updated 3 weeks ago
- a simple c++ inference engine for gpt based architecture☆40Dec 10, 2025Updated 8 months ago
- ☆15Jul 25, 2025Updated last year
- ☆50Jul 14, 2026Updated 3 weeks ago
- ☆17Dec 28, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Flash Attention from scratch, tiled CUDA forward kernel, online softmax with running max and correction factor, recomputation trick in ba…☆18Mar 6, 2026Updated 5 months ago
- ☆26Dec 5, 2022Updated 3 years ago
- ☆19Apr 26, 2026Updated 3 months ago
- Optimizing diffusion for production-ready speeds☆40Jan 10, 2026Updated 7 months ago
- Lets build a Deep Learning Framework!☆28Mar 12, 2026Updated 4 months ago
- raft based distributed kv store☆29Jan 6, 2026Updated 7 months ago
- ☆13Apr 7, 2025Updated last year
- 化学式自动配平计算器☆11Jul 28, 2022Updated 4 years ago
- "JABAS: Joint Adaptive Batching and Automatic Scaling for DNN Training on Heterogeneous GPUs" (EuroSys '25)☆16Apr 7, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official codebase for the MLSys 2026 paper "IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference". It enables hi…☆20May 29, 2026Updated 2 months ago
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- serving a torch model using Celery, Redis and RabbitMQ to serve users asynchronously☆25Jan 21, 2024Updated 2 years ago
- egraph <-> json☆17Dec 29, 2025Updated 7 months ago
- Hacks for PyTorch☆19Apr 18, 2023Updated 3 years ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆239Jun 29, 2026Updated last month
- Simple intermediate representation language for learning and research.☆22Mar 27, 2020Updated 6 years ago
- Official repository for the paper DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines☆19Dec 8, 2023Updated 2 years ago
- Official resporitory for "IPDPS' 24 QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices".☆20Feb 23, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- fastcv is a CUDA rewrite of the opencv filters with python bindings☆73Oct 10, 2025Updated 10 months ago
- Demo code for CVPR2023 paper "Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision Transformers"☆15Jul 4, 2023Updated 3 years ago
- Using e-graphs to synthesize netlists from boolean logic.☆14Jul 26, 2023Updated 3 years ago
- A simple plugin to automatically restart the server after the server crashed☆10Aug 22, 2025Updated 11 months ago
- Async utilities for Golang☆16Nov 22, 2020Updated 5 years ago
- Generates text with diffusion models. Reproduction of the Continous Diffusion for Categorical Data paper by Deepmind☆18Dec 9, 2024Updated last year
- AI based singing voice synthesis database generator☆13Aug 12, 2022Updated 3 years ago
- Hi-Speed DNN Training with Espresso: Unleashing the Full Potential of Gradient Compression with Near-Optimal Usage Strategies (EuroSys '2…☆15Sep 21, 2023Updated 2 years ago
- Made with LaTex. NENU's recommendation letter template.☆12May 26, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Simple RAM benchmark for Linux.☆12Aug 4, 2021Updated 5 years ago
- ☆14Jan 18, 2023Updated 3 years ago
- Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling☆13Mar 7, 2024Updated 2 years ago
- I love reinforcement learning.☆12Jan 15, 2025Updated last year
- An expense tracking application that embodiesminimalist design philosophy with clean typography, subtle shadows, and elegant micro-intera…☆15Aug 31, 2025Updated 11 months ago
- ☆22Dec 15, 2023Updated 2 years ago
- Interactive version of the CuTe layout paper☆57Apr 14, 2026Updated 3 months ago