Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V
☆182Jun 30, 2026Updated last month
Alternatives and similar repositories for flash-attention-v100
Users that are interested in flash-attention-v100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆18Feb 9, 2026Updated 6 months ago
- forked from vllm-project/flash-attention☆59May 9, 2026Updated 3 months ago
- vLLM fork for Tesla V100 (SM70) with AWQ 4-bit support, CUDA 12.8 build flow, and validated Qwen3.5 27B/35B deployment on multi-GPU V…☆586Updated this week
- vLLM fork for Tesla V100 (SM70) — extends 1CatAI's AWQ support and adds GGUF support☆20Jun 20, 2026Updated last month
- 1CatV2 with TileLANG written FA-v100 and many goodies☆18Jul 15, 2026Updated 3 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆21Jul 2, 2026Updated last month
- ☆78Feb 19, 2024Updated 2 years ago
- NVIDIA Tesla V100 显卡驱动自动化管理套件☆26Oct 9, 2025Updated 10 months ago
- vllm混合推理扩展插件,支持多NUMA混合推理,单卡推理Qwen3-Next模型可达1000+ prefill☆34Nov 7, 2025Updated 9 months ago
- ☆35Mar 26, 2025Updated last year
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 6 months ago
- Flash Attention 2 implementation for Turing GPUs☆120Mar 23, 2026Updated 4 months ago
- NVIDIA Linux open GPU with P2P support☆392Jul 13, 2026Updated 3 weeks ago
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆105Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- This suite of nodes unlocks high-performance parallel processing in ComfyUI by utilizing **Model Replication**. Unlike standard offloadin…☆60Feb 24, 2026Updated 5 months ago
- A ComfyUI custom node implementation of TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows.☆44Mar 6, 2026Updated 5 months ago
- ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆301Updated this week
- Ultra-Sortformer for Scalable Speaker Diarization☆27Apr 9, 2026Updated 4 months ago
- Force any Model, VAE, or CLIP to any GPU or CPU—and keep it there (or don't).☆33Feb 6, 2026Updated 6 months ago
- Give OpenClaw the power to control your desktop — UI-TARS VLM-driven GUI automation☆18Feb 17, 2026Updated 5 months ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆37Jul 22, 2026Updated 2 weeks ago
- RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs☆20Feb 8, 2026Updated 6 months ago
- WanImageToVideo ComfyUI node, with Tiled VAE☆16Oct 22, 2025Updated 9 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 6 months ago
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆222Updated this week
- The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode…☆483Jul 13, 2026Updated 3 weeks ago
- A plugin code to have a simple face instead of the Xiaozhi AI firmware☆27Aug 1, 2026Updated last week
- ComfyUI-Bagel is now available in ComfyUI, BAGEL is an open‑source multimodal foundation model with 7B active parameters (14B total) trai…☆29May 28, 2025Updated last year
- Processing of astronomical images☆15Jul 17, 2026Updated 3 weeks ago
- 4-bit Shampoo for Memory-Efficient Network Training (NeurIPS 2024)☆13Feb 13, 2025Updated last year
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆852Updated this week
- OpenCode plugins that help AI agents understand their identity across multi-agent sessions☆28Apr 20, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆77Jul 14, 2026Updated 3 weeks ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, ex…☆25Aug 3, 2026Updated last week
- 谷歌INCEPTION-RESNET-V2迁移学习实现图像二分类判断图像是否生病☆17Apr 23, 2018Updated 8 years ago
- ☆69Nov 11, 2025Updated 9 months ago
- ZLUDA PTX test suite☆23Feb 17, 2026Updated 5 months ago
- Automaton & Cognition☆17Apr 14, 2024Updated 2 years ago
- ☆55Updated this week