a minimal paged attention implementation
☆20Jan 30, 2026Updated 7 months ago
Alternatives and similar repositories for nano-paged-attention
Users that are interested in nano-paged-attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- writing really fast kernels☆19Jul 15, 2026Updated last month
- ☆15Jul 25, 2025Updated last year
- ☆132Dec 9, 2025Updated 8 months ago
- a LLM inference engine to run on consumer hardware☆47Apr 15, 2026Updated 4 months ago
- ☆51Aug 17, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Flash Attention from scratch, tiled CUDA forward kernel, online softmax with running max and correction factor, recomputation trick in ba…☆18Mar 6, 2026Updated 5 months ago
- ☆26Dec 5, 2022Updated 3 years ago
- ☆12Jun 14, 2024Updated 2 years ago
- ☆27Jun 7, 2026Updated 2 months ago
- ☆20Apr 26, 2026Updated 4 months ago
- Optimizing diffusion for production-ready speeds☆40Jan 10, 2026Updated 7 months ago
- Lets build a Deep Learning Framework!☆28Mar 12, 2026Updated 5 months ago
- MathNet: A Data-Centric Approach, Dataset and Benchmark Model to Advance Mathematical Expression Recognition☆10Mar 19, 2025Updated last year
- A high-performance job scheduler designed for microsecond latency (<5us P50 latency) and massive concurrency (highest throughput : 2M+ jo…☆30Jan 11, 2026Updated 7 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆13Apr 7, 2025Updated last year
- Official codebase for the MLSys 2026 paper "IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference". It enables hi…☆21May 29, 2026Updated 3 months ago
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- serving a torch model using Celery, Redis and RabbitMQ to serve users asynchronously☆25Jan 21, 2024Updated 2 years ago
- egraph <-> json☆17Dec 29, 2025Updated 8 months ago
- It contains Data Augmentaion, Strided convolution, Batch Normalization, Leaky Relu, Global Average pooling, L2 Regularization, learning …☆12Jun 3, 2018Updated 8 years ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆240Jun 29, 2026Updated 2 months ago
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 4 months ago
- Simple intermediate representation language for learning and research.☆22Mar 27, 2020Updated 6 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official repository for the paper DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines☆19Dec 8, 2023Updated 2 years ago
- Official resporitory for "IPDPS' 24 QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices".☆20Feb 23, 2024Updated 2 years ago
- Demo code for CVPR2023 paper "Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision Transformers"☆15Jul 4, 2023Updated 3 years ago
- Online Multiplayer UNO Game☆23Sep 4, 2023Updated 2 years ago
- Using e-graphs to synthesize netlists from boolean logic.☆14Jul 26, 2023Updated 3 years ago
- Perceptron-based branch predictor written in C++☆14Dec 14, 2016Updated 9 years ago
- A visual representation of Dijkstra's Algorithm using Libgdx.☆16Dec 24, 2021Updated 4 years ago
- AI based singing voice synthesis database generator☆13Aug 12, 2022Updated 4 years ago
- Generates text with diffusion models. Reproduction of the Continous Diffusion for Categorical Data paper by Deepmind☆19Dec 9, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Hi-Speed DNN Training with Espresso: Unleashing the Full Potential of Gradient Compression with Near-Optimal Usage Strategies (EuroSys '2…☆15Sep 21, 2023Updated 2 years ago
- List of all the resources I used during 10 days of Statistics and Data Preprocessing.☆16Jan 4, 2021Updated 5 years ago
- Made with LaTex. NENU's recommendation letter template.☆12May 26, 2024Updated 2 years ago
- Simple RAM benchmark for Linux.☆12Aug 4, 2021Updated 5 years ago
- ☆14Jan 18, 2023Updated 3 years ago
- XoRL☆43Updated this week
- ☆22Dec 15, 2023Updated 2 years ago