High-performance FlashAttention-2 for AMD, Intel, and Apple GPUs. Drop-in replacement for PyTorch SDPA. Triton backend for ROCm (MI300X, RDNA3), Vulkan backend for consumer GPUs. No CUDA required.
☆158Jan 27, 2026Updated 5 months ago
Alternatives and similar repositories for Aule-Attention
Users that are interested in Aule-Attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆11Feb 20, 2025Updated last year
- WanImageToVideo ComfyUI node, with Tiled VAE☆16Oct 22, 2025Updated 8 months ago
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 10 months ago
- A sleek, customizable interface for managing LLMs with responsive design and easy agent personalization.☆18Aug 30, 2024Updated last year
- ☆25Feb 10, 2026Updated 5 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Simple monkeypatch to boost AMD Navi 3 GPUs☆51Apr 21, 2025Updated last year
- ☆32Updated this week
- ☆40Apr 29, 2024Updated 2 years ago
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,160Updated this week
- Text-generation-webui oneclick UI + API☆12Dec 27, 2024Updated last year
- Local-first desktop AI workbench for roleplay, multi-character chat, long-form writing, RAG, MCP tools, plugins, and local models.☆104Updated this week
- ☆24Jan 22, 2025Updated last year
- Tcurtsni: Reverse Instruction Chat, ever wonder what your LLM wants to ask you?☆23Jun 25, 2024Updated 2 years ago
- A vector similarity search engine for humans🥳☆18Oct 30, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆19Jul 2, 2026Updated 2 weeks ago
- Privacy-first agentic framework with powerful reasoning & task automation capabilities. Natively distributed and fully ISO 27XXX complian…☆70Apr 1, 2025Updated last year
- ☆23Sep 27, 2024Updated last year
- ☆15Jan 27, 2023Updated 3 years ago
- ☆92Dec 16, 2025Updated 7 months ago
- Home Made Diffusion Models☆199Dec 9, 2025Updated 7 months ago
- ScribePal is an Open Source intelligent browser extension that leverages AI to empower your web experience by providing contextual insigh…☆22Apr 6, 2026Updated 3 months ago
- Luth is a state-of-the-art series of fine-tuned LLMs for French☆46Oct 12, 2025Updated 9 months ago
- ComfyUI-ThinkSound is now available in ComfyUI, ThinkSound is a unified Any2Audio generation framework with flow matching guided by Chain…☆29Jul 12, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆23Mar 21, 2026Updated 4 months ago
- AI Based "Happiness Optimizer"☆12Oct 20, 2024Updated last year
- REAP: Router-weighted Expert Activation Pruning for SMoE compression☆447Apr 17, 2026Updated 3 months ago
- MagicNodes, it's a plug-and-play multi-pass "render-machine" for SD/SDXL models. Simple one-node start, expert-grade results. Core is ZeR…☆91May 16, 2026Updated 2 months ago
- Run ComfyUI on Modal with auto-scaling, GPU snapshots, and easy model management.☆34Jun 30, 2026Updated 3 weeks ago
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆34Jul 14, 2026Updated last week
- [ICML 2026 Spotlight] Official implementation of TetraJet-v2: Accurate NVFP4 Training for LLMs, with fully-NVFP4 linear layer with unbias…☆17Jul 3, 2026Updated 2 weeks ago
- The most powerful and modular stable diffusion GUI, api and backend with a graph/nodes interface. Now ZLUDA enhanced for better AMD GPU p…☆931Updated this week
- Official Implementation of Knowledge Flow Prompting☆35Oct 20, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Distribute and run LLMs with a single file.☆25May 13, 2025Updated last year
- ☆97Apr 26, 2025Updated last year
- (ECCV 2026): Official code for Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models☆18Jul 9, 2026Updated last week
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆24Nov 26, 2025Updated 7 months ago
- ☆22Jul 25, 2023Updated 2 years ago
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,056Updated this week
- Lower Precision Floating Point Operations☆83Feb 22, 2026Updated 4 months ago