Flash Attention 2 implementation for Turing GPUs
☆116Mar 23, 2026Updated 3 months ago
Alternatives and similar repositories for flash-attention-turing
Users that are interested in flash-attention-turing are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆35Mar 26, 2025Updated last year
- Cross-platform FlashAttention-2 Triton implementation for Turing+ GPUs with custom configuration mode☆26Jan 12, 2026Updated 6 months ago
- Sage attention for turning.☆70Dec 29, 2025Updated 6 months ago
- Custom tool set mostly for Hunyuan Video, but includes some WAN and FramePack Video nodes.☆18Feb 16, 2026Updated 5 months ago
- ComfyUI nodes for TencentARC Pixal3D image-to-3D generation☆26Jun 1, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆78Feb 19, 2024Updated 2 years ago
- NES emulator written in pure FreeBASIC with love by Blyss Sarania and Gavin Schulte(Nobbs66).☆21Oct 29, 2025Updated 8 months ago
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆175Jun 30, 2026Updated 3 weeks ago
- macrogpt大模型全量预训练(1b3,32层), 多卡deepspeed/单卡adafactor☆15Nov 30, 2023Updated 2 years ago
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode…☆431Jul 13, 2026Updated last week
- ☆45Nov 1, 2025Updated 8 months ago
- Standalone Flash Attention v2 kernel without libtorch dependency☆113Sep 10, 2024Updated last year
- async inference for machine learning model☆26Sep 21, 2022Updated 3 years ago
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆25Sep 1, 2025Updated 10 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ComfyUI custom node for Qwen3.5-9B unified multimodal model☆38Mar 13, 2026Updated 4 months ago
- ☆12Mar 21, 2024Updated 2 years ago
- ☆26Jun 11, 2025Updated last year
- minimal C implementation of speculative decoding based on llama2.c☆30Jul 15, 2024Updated 2 years ago
- 🎭 character card editor online.☆36Apr 13, 2026Updated 3 months ago
- Implementation of multi-level Contrastive Predictive Coding (CPC) methods☆20Jan 12, 2023Updated 3 years ago
- some helper script to tagging with DeepDanbooru and BLIP☆31Nov 27, 2022Updated 3 years ago
- Raspberry Pi 5 with Arch Linux ARM and encrypted root☆14Jul 20, 2024Updated 2 years ago
- Javascripts Deobfuscator. Used to debug obfuscated JS from obfuscator.io and other obfuscate tools.☆11Oct 25, 2023Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆15Feb 9, 2026Updated 5 months ago
- LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features …☆386Updated this week
- ☆17Apr 4, 2026Updated 3 months ago
- Redirect requests of current origin to another domain with Service Worker.☆13Dec 13, 2022Updated 3 years ago
- ☆11Apr 13, 2019Updated 7 years ago
- Syntax Coloring for Freebasic in VScode☆13Mar 19, 2024Updated 2 years ago
- Tries to UI development. Clone of https://www.perplexity.ai/☆11Sep 30, 2023Updated 2 years ago
- [ICML 2025] Adaptive Self-improvement LLM Agentic System for ML Library Development☆17Jan 6, 2026Updated 6 months ago
- ☆15Dec 21, 2025Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Communicate undetected in plain sight using zero width obfuscation.☆15Nov 5, 2021Updated 4 years ago
- Generic card scanner meant for card games with similar card formats as MTG & Pokemon. Takes an image, scans for a card, scans card for na…☆14Sep 7, 2020Updated 5 years ago
- 基于EventLoop和多线程的morden cpp 的linux网络库☆11Apr 5, 2020Updated 6 years ago
- GPU accelerated Perlin Noise in python☆11Oct 23, 2020Updated 5 years ago
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆286Jul 8, 2026Updated last week
- 一个基于 Playwright 和AI过滤分析的闲鱼多任务实时监控与智能分析工具,配备了功能完善的 Web 管理界面。☆16Jan 13, 2026Updated 6 months ago
- Simulation backend for KOS☆15May 6, 2025Updated last year