We aim to redefine Data Parallel libraries portabiliy, performance, programability and maintainability, by using C++ standard features, instead of creating new compilers.
☆58Sep 13, 2026Updated last week
Alternatives and similar repositories for FusedKernelLibrary
Users that are interested in FusedKernelLibrary are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A faster implementation of OpenCV-CUDA that uses OpenCV objects, and more!☆56Sep 17, 2026Updated last week
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆82Feb 18, 2026Updated 7 months ago
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Code☆16Mar 19, 2023Updated 3 years ago
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆25Apr 10, 2026Updated 5 months ago
- Open-source transpiler for CUDA Tile (13.1) migration☆19Dec 9, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆20Apr 26, 2026Updated 4 months ago
- A source-to-source compiler for optimizing CUDA dynamic parallelism by aggregating launches☆15Jun 21, 2019Updated 7 years ago
- A fuzzer for ML compilers☆47Aug 7, 2026Updated last month
- ☆17Oct 21, 2020Updated 5 years ago
- HiCOPS: Computational framework for peptide identification from MS data through accelerated database search☆10Mar 24, 2023Updated 3 years ago
- An auxiliary project analysis of the characteristics of KV in DiT Attention.☆34Nov 29, 2024Updated last year
- diffusers with search engine☆12Aug 12, 2026Updated last month
- Spatial runtime for real-time pose estimation, mapping, and agent interaction.☆40Updated this week
- Comb is a communication performance benchmarking tool.☆25Feb 27, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Framework to reduce autotune overhead to zero for well known deployments.☆101Sep 19, 2025Updated last year
- A simple environment for writing and experimenting with hand-written CUDA PTX kernels.☆19Sep 11, 2025Updated last year
- EleutherAI ML Performance reading group repository (slides, meeting recordings, annotated papers)☆37Mar 20, 2026Updated 6 months ago
- ☆19May 18, 2026Updated 4 months ago
- Artifacts of EVT ASPLOS'24☆29Mar 6, 2024Updated 2 years ago
- CLIP and SigLIP models optimized with TensorRT with a Transformers-like API☆37Sep 29, 2024Updated last year
- High-performance LLM operator library built on TileLang.☆190Updated this week
- Zero-copy multimodal vector DB with CUDA and CLIP/SigLIP☆66May 6, 2025Updated last year
- Kernel Library for Large Headdim Attention (64~1024, BF16/FP8/FP4), 1.5x~15x↑ vs PyTorch SDPA.☆334Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A Triton-only attention backend for vLLM☆28Jul 14, 2026Updated 2 months ago
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆77Aug 8, 2026Updated last month
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- ☆27Aug 28, 2025Updated last year
- Tensor Algebra for many-body methods☆18Sep 2, 2026Updated 3 weeks ago
- ☆45Sep 8, 2025Updated last year
- Code to handle I/O from/to data in CASA format☆11Sep 17, 2026Updated last week
- A dynamic GPU memory allocator, suitable for warp synchronized scenarios.☆11Aug 20, 2019Updated 7 years ago
- Scalable GPU Kernel Fission/Fusion Transformation for Memory-Bound Kernels☆14Aug 26, 2015Updated 11 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆35Apr 2, 2025Updated last year
- flash attention 优化日志☆34Aug 28, 2026Updated 3 weeks ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo☆27Sep 16, 2026Updated last week
- OpenMM plugin to interface with XTB☆20Nov 5, 2025Updated 10 months ago
- A Library for intra-GPU/Inter-SM parallelsim☆13Aug 7, 2026Updated last month
- Multi-stage LLM agent pipeline for optimizing Triton kernels on Intel XPU — from analysis to autotuning.☆22Sep 11, 2026Updated last week
- DXT Explorer is an interactive web-based log analysis tool for Darshan DXT logs.☆18Feb 19, 2026Updated 7 months ago