Fast and memory-efficient exact attention
☆34Dec 2, 2024Updated last year
Alternatives and similar repositories for flash-attention-3
Users that are interested in flash-attention-3 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Implementation of Infinite-Resolution Integral Noise Warping for Diffusion Models [ICLR 2025]☆16Mar 15, 2025Updated last year
- Block Diffusion Trainer☆15Jul 10, 2025Updated last year
- Basic world models☆32Oct 30, 2025Updated 8 months ago
- SGLang Kernel Wheel Index☆24Updated this week
- ☆135May 29, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆24Jun 18, 2024Updated 2 years ago
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspace☆19Oct 21, 2024Updated last year
- ☆27May 3, 2024Updated 2 years ago
- ☆13Jun 16, 2026Updated last month
- ☆44Jun 19, 2024Updated 2 years ago
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"☆18Mar 15, 2024Updated 2 years ago
- CUDA SGEMM optimization note☆15Oct 31, 2023Updated 2 years ago
- A curated reading list on harness engineering for recursive self-improvement of LLM agents (EN/ZH).☆16Jul 9, 2026Updated 2 weeks ago
- Hardware Division Units☆10Jul 17, 2014Updated 12 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR'25] Official repository of paper: Ranking-aware adapter for text-driven image ordering with CLIP☆16Apr 17, 2025Updated last year
- This repository contains code for the MicroAdam paper.☆21Dec 14, 2024Updated last year
- NetHCF: Enabling Line-rate and Adaptive Spoofed IP Traffic Filtering☆13Mar 17, 2022Updated 4 years ago
- RADIX-4 SRT division☆12Oct 31, 2019Updated 6 years ago
- 国自然基金Latex模板☆11Mar 12, 2023Updated 3 years ago
- A CUDA kernel for NHWC GroupNorm for PyTorch☆23Nov 15, 2024Updated last year
- In this project, we propose to study Vision Transformers trained using the Barlow Twins self-supervised method, and compare the results w…☆17Oct 3, 2023Updated 2 years ago
- Acclaim: Adaptive Memory Reclaim to Improve User Experience in Android Systems [ATC '20]☆16Aug 1, 2020Updated 5 years ago
- Make-A-Video Latent Diffusion Model☆19Nov 15, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- This is the official implementation for "Deep Magnification-Flexible Upsampling over3D Point Clouds".☆13Dec 2, 2021Updated 4 years ago
- [AAAI23] Parametric Surface Constrained Upsampler Network for Point Cloud☆17Jul 12, 2024Updated 2 years ago
- ☆13Jun 4, 2024Updated 2 years ago
- Immix GC for LLVM based languages☆17Apr 2, 2025Updated last year
- JAX bindings for the flash-attention3 kernels☆23Jan 2, 2026Updated 6 months ago
- ☆21Jun 26, 2023Updated 3 years ago
- An Agile RISC-V SoC Design Framework with in-order cores, out-of-order cores, accelerators, and more☆12May 29, 2026Updated last month
- Evaluate state-of-the-art GPU joins☆14Nov 29, 2023Updated 2 years ago
- Basic floating-point components for RISC-V processors☆12Aug 13, 2017Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Triton implementation of Flash Attention2.0☆54Jul 31, 2023Updated 2 years ago
- Optimizing Tensor Computation Graphs with Equality Saturation and Monte Carlo Tree Search☆15Aug 9, 2024Updated last year
- A implement of run-length encoding for Pytorch tensor using CUDA☆14Apr 7, 2021Updated 5 years ago
- Advanced Video Graph RAG using SAM2,CLIP,BLIP,Qwen2-VL,YOLO-World ,Neo4j, WebGPU, local LLM☆14Nov 25, 2024Updated last year
- RADLADS training code☆46May 7, 2025Updated last year
- Jina VDR is a multilingual, multi-domain benchmark for visual document retrieval☆38Aug 4, 2025Updated 11 months ago
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆15Jul 4, 2026Updated 2 weeks ago