High-performance FlashAttention-2 for AMD, Intel, and Apple GPUs. Drop-in replacement for PyTorch SDPA. Triton backend for ROCm (MI300X, RDNA3), Vulkan backend for consumer GPUs. No CUDA required.
☆160Jan 27, 2026Updated 7 months ago
Alternatives and similar repositories for Aule-Attention
Users that are interested in Aule-Attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆11Feb 20, 2025Updated last year
- WanImageToVideo ComfyUI node, with Tiled VAE☆16Oct 22, 2025Updated 10 months ago
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated last year
- A sleek, customizable interface for managing LLMs with responsive design and easy agent personalization.☆19Aug 30, 2024Updated 2 years ago
- ☆25Feb 10, 2026Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Simple monkeypatch to boost AMD Navi 3 GPUs☆51Apr 21, 2025Updated last year
- ☆35Aug 9, 2026Updated last month
- Text-generation-webui oneclick UI + API☆12Dec 27, 2024Updated last year
- A module for trainable encoder/decoder filterbanks with auditory bias.☆17Feb 17, 2026Updated 7 months ago
- Local-first desktop AI workbench for roleplay, multi-character chat, long-form writing, RAG, MCP tools, plugins, and local models.☆134Updated this week
- ☆24Jan 22, 2025Updated last year
- Tcurtsni: Reverse Instruction Chat, ever wonder what your LLM wants to ask you?☆23Jun 25, 2024Updated 2 years ago
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆24Jul 2, 2026Updated 2 months ago
- a simple Flash Attention v2 implementation with ROCM (RDNA3 GPU, roc wmma), mainly used for stable diffusion(ComfyUI) in Windows ZLUDA en…☆54Aug 25, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Privacy-first agentic framework with powerful reasoning & task automation capabilities. Natively distributed and fully ISO 27XXX complian…☆69Apr 1, 2025Updated last year
- ☆23Sep 27, 2024Updated last year
- ☆15Jan 27, 2023Updated 3 years ago
- Efficient non-uniform quantization with GPTQ for GGUF☆66Sep 17, 2025Updated last year
- Home Made Diffusion Models☆201Dec 9, 2025Updated 9 months ago
- ScribePal is an Open Source intelligent browser extension that leverages AI to empower your web experience by providing contextual insigh…☆22Sep 13, 2026Updated last week
- Luth is a state-of-the-art series of fine-tuned LLMs for French☆47Oct 12, 2025Updated 11 months ago
- ComfyUI-ThinkSound is now available in ComfyUI, ThinkSound is a unified Any2Audio generation framework with flow matching guided by Chain…☆29Jul 12, 2025Updated last year
- ☆17Aug 31, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆13Jul 19, 2023Updated 3 years ago
- AI Based "Happiness Optimizer"☆12Oct 20, 2024Updated last year
- ☆31Aug 27, 2026Updated 3 weeks ago
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆122Jun 10, 2026Updated 3 months ago
- A PyTorch implementation of gradient-free optimization for directly optimizing NDCG (Normalized Discounted Cumulative Gain) in neural inf…☆19Dec 21, 2025Updated 8 months ago
- ☆83Apr 14, 2026Updated 5 months ago
- RadialAttention in ComfyUI native workflow☆121Dec 19, 2025Updated 9 months ago
- Run ComfyUI on Modal with auto-scaling, GPU snapshots, and easy model management. Try image, video generation via ComfyUI on Modal.☆44Sep 6, 2026Updated 2 weeks ago
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆37Sep 13, 2026Updated last week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Protocol for Augmented Memory of Project Artifacts (MCP compatible) - extended☆23Jan 24, 2026Updated 7 months ago
- Distribute and run LLMs with a single file.☆25May 13, 2025Updated last year
- [ICML 2026 Spotlight] Official implementation of TetraJet-v2: Accurate NVFP4 Training for LLMs, with fully-NVFP4 linear layer with unbias…☆17Jul 3, 2026Updated 2 months ago
- Apple's Neural Engine - bare metal access☆45Jul 2, 2026Updated 2 months ago
- ☆97Apr 26, 2025Updated last year
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆25Nov 26, 2025Updated 9 months ago
- Lower Precision Floating Point Operations☆86Feb 22, 2026Updated 6 months ago