We aim to redefine Data Parallel libraries portabiliy, performance, programability and maintainability, by using C++ standard features, instead of creating new compilers.
☆55Sep 1, 2026Updated this week
Alternatives and similar repositories for FusedKernelLibrary
Users that are interested in FusedKernelLibrary are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A faster implementation of OpenCV-CUDA that uses OpenCV objects, and more!☆55Aug 27, 2026Updated last week
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 6 months ago
- Effective transpose on Hopper GPU☆29Sep 6, 2025Updated 11 months ago
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆25Apr 10, 2026Updated 4 months ago
- Open-source transpiler for CUDA Tile (13.1) migration☆19Dec 9, 2025Updated 8 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- PIRA - Automatic Instrumentation Refinement☆18Mar 28, 2024Updated 2 years ago
- A fuzzer for ML compilers☆47Aug 7, 2026Updated 3 weeks ago
- ☆17Oct 21, 2020Updated 5 years ago
- HiCOPS: Computational framework for peptide identification from MS data through accelerated database search☆10Mar 24, 2023Updated 3 years ago
- Spatial runtime for real-time pose estimation, mapping, and agent interaction.☆36Updated this week
- An auxiliary project analysis of the characteristics of KV in DiT Attention.☆34Nov 29, 2024Updated last year
- diffusers with search engine☆12Aug 12, 2026Updated 3 weeks ago
- An agent for CUDA compute-communication kernel co-design☆37May 7, 2026Updated 3 months ago
- JAX Multi-Agent RL, Neuro-Evolution, and A-Life Library☆15Oct 12, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Awesome code, projects, books, etc. related to CUDA☆41Aug 9, 2026Updated 3 weeks ago
- Comb is a communication performance benchmarking tool.☆25Feb 27, 2023Updated 3 years ago
- A simple environment for writing and experimenting with hand-written CUDA PTX kernels.☆19Sep 11, 2025Updated 11 months ago
- Framework to reduce autotune overhead to zero for well known deployments.☆101Sep 19, 2025Updated 11 months ago
- EleutherAI ML Performance reading group repository (slides, meeting recordings, annotated papers)☆37Mar 20, 2026Updated 5 months ago
- Artifacts of EVT ASPLOS'24☆29Mar 6, 2024Updated 2 years ago
- Prototype for a SPIR-V assembler and dissasembler. It provides a composable Java interface for generating SPIR-V code at runtime.☆16Oct 31, 2025Updated 10 months ago
- CLIP and SigLIP models optimized with TensorRT with a Transformers-like API☆36Sep 29, 2024Updated last year
- High-performance LLM operator library built on TileLang.☆177Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Adelic p-adic Dark Matter☆15Apr 25, 2026Updated 4 months ago
- Zero-copy multimodal vector DB with CUDA and CLIP/SigLIP☆65May 6, 2025Updated last year
- Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.☆331Updated this week
- A Triton-only attention backend for vLLM☆28Jul 14, 2026Updated last month
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆76Aug 8, 2026Updated 3 weeks ago
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- ☆27Aug 28, 2025Updated last year
- Tensor Algebra for many-body methods☆18Updated this week
- ☆45Sep 8, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆22Updated this week
- A dynamic GPU memory allocator, suitable for warp synchronized scenarios.☆11Aug 20, 2019Updated 7 years ago
- Scalable GPU Kernel Fission/Fusion Transformation for Memory-Bound Kernels☆14Aug 26, 2015Updated 11 years ago
- ☆51Jan 18, 2024Updated 2 years ago
- ☆35Apr 2, 2025Updated last year
- flash attention 优化日志☆34Aug 28, 2026Updated last week
- [DEPRECATED] Moved to ROCm/rocm-libraries repo☆27Aug 18, 2026Updated 2 weeks ago