Flash Attention 2 implementation for Turing GPUs
☆126Sep 14, 2026Updated last week
Alternatives and similar repositories for flash-attention-turing
Users that are interested in flash-attention-turing are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cross-platform FlashAttention-2 Triton implementation for Turing+ GPUs with custom configuration mode☆33Jan 12, 2026Updated 8 months ago
- ☆20Jul 31, 2026Updated last month
- ComfyUI nodes for TencentARC Pixal3D image-to-3D generation☆29Jun 1, 2026Updated 3 months ago
- ☆78Feb 19, 2024Updated 2 years ago
- Scripts for fbfrog-based FreeBASIC bindings☆17Mar 16, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆207Jun 30, 2026Updated 2 months ago
- AI21 Typescript SDK☆13Dec 18, 2025Updated 9 months ago
- macrogpt大模型全量预训练(1b3,32层), 多卡deepspeed/单卡adafactor☆15Nov 30, 2023Updated 2 years ago
- Fix Nov. 2024☆19Sep 16, 2025Updated last year
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 200+ tok/s single-reques…☆906Updated this week
- triton for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆48Dec 8, 2025Updated 9 months ago
- ComfyUI node to run grounding models☆49Sep 14, 2026Updated last week
- ☆45Nov 1, 2025Updated 10 months ago
- Standalone Flash Attention v2 kernel without libtorch dependency☆113Sep 10, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- async inference for machine learning model☆26Sep 21, 2022Updated 4 years ago
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆26Sep 1, 2025Updated last year
- ☆47Jun 18, 2026Updated 3 months ago
- ComfyUI custom node for Qwen3.5-9B unified multimodal model☆40Mar 13, 2026Updated 6 months ago
- ☆12Mar 21, 2024Updated 2 years ago
- ComfyUI Wrapper for Microsoft Trellis.2 - Native and Compact Structured Latents for 3D Generation☆104Aug 16, 2026Updated last month
- a simple wrapper of fooocus prompt expansion engine in stable-diffusion-webui☆23Jun 9, 2024Updated 2 years ago
- Enclosed chamber-type Delta 3d printer design☆11Aug 28, 2023Updated 3 years ago
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆457Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 네이버 블로그 원본 이미지 크롤러☆10Feb 18, 2023Updated 3 years ago
- ☆16Jun 25, 2026Updated 2 months ago
- Decensoring Hentai☆13Sep 19, 2022Updated 4 years ago
- PPE detection of helmets(construction) using Nvidia Deepstream. Model trained using Nvidia TLT.☆11Jun 27, 2021Updated 5 years ago
- Fast and memory-efficient exact attention☆23Jun 26, 2026Updated 2 months ago
- [ICML 2025] Adaptive Self-improvement LLM Agentic System for ML Library Development☆17Jan 6, 2026Updated 8 months ago
- ☆14Sep 4, 2024Updated 2 years ago
- 基于EventLoop和多线程的morden cpp 的linux网络库☆11Apr 5, 2020Updated 6 years ago
- GPU accelerated Perlin Noise in python☆11Oct 23, 2020Updated 5 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆17Sep 1, 2024Updated 2 years ago
- Simulation backend for KOS☆15May 6, 2025Updated last year
- Tabby, a Self Hosted way to save and manage Bookmarks☆13Nov 21, 2020Updated 5 years ago
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆70May 4, 2025Updated last year
- Julia implementation of flash-attention operation for neural networks.☆11May 31, 2023Updated 3 years ago
- A simple tool to easily use Montreal Forced Aligner. Also provide alignment(TextGrid) retrieved from ESD.☆45May 25, 2023Updated 3 years ago
- 基于 OpenCV dnn 实现的 MTCNN 人脸检测器☆13Mar 31, 2019Updated 7 years ago