β21Mar 17, 2026Updated 5 months ago
Alternatives and similar repositories for redfuser
Users that are interested in redfuser are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β24Sep 1, 2025Updated 11 months ago
- πAutomatically Update LLM inference systems Papers Daily using Github Actions (Update Every 12th hours)β12Updated this week
- A DL compiler fuzzerβ15Nov 1, 2024Updated last year
- A Triton-only attention backend for vLLMβ28Jul 14, 2026Updated last month
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitiveβ76Aug 8, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β19Mar 4, 2025Updated last year
- β16Jan 24, 2024Updated 2 years ago
- DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scalingβ38Aug 19, 2026Updated last week
- β27Feb 20, 2024Updated 2 years ago
- OSDI 2023 Welder, deeplearning compilerβ35Nov 24, 2023Updated 2 years ago
- Autonomous CUDA kernel optimization agent with structured task specs and per-config scoringβ17Jun 17, 2026Updated 2 months ago
- β26Jun 10, 2026Updated 2 months ago
- Verified graph rewriting (for dataflow circuits).β27Jul 15, 2026Updated last month
- β30Apr 7, 2025Updated last year
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Multiple 1-stencil implementations using nvidia cuda.β12Dec 2, 2017Updated 8 years ago
- Simple python library for generating your own perfetto traces for your application. Can be used for both app instrumentation and custom β¦β27Jun 22, 2025Updated last year
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batchβ¦β135Jun 29, 2026Updated 2 months ago
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.β19Jun 12, 2026Updated 2 months ago
- β27Oct 1, 2025Updated 10 months ago
- Discovery of Structured Parallelism In Sequential and Parallel Codeβ10Feb 13, 2021Updated 5 years ago
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Codeβ16Mar 19, 2023Updated 3 years ago
- β33Jul 17, 2024Updated 2 years ago
- β10Mar 8, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- β162Aug 8, 2026Updated 3 weeks ago
- GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weavingβ21Jul 30, 2025Updated last year
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and minβ¦β10Aug 13, 2024Updated 2 years ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.β117Aug 4, 2026Updated 3 weeks ago
- β200Aug 9, 2026Updated 2 weeks ago
- Website for CSE 234, Winter 2025β16Mar 24, 2025Updated last year
- SIMPLER MAGIC: Synthesis and In-memory MaPping of Logic Execution in a single Row for Memristor Aided loGICβ13Dec 5, 2019Updated 6 years ago
- MIT 6.172 Performance Engineering of Software Systemsβ17Dec 30, 2021Updated 4 years ago
- Share your GPU without MIG or MPSβ50Jan 27, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β46Oct 15, 2025Updated 10 months ago
- A collection of out-of-tree extensions for the Triton language and compilerβ34Updated this week
- A new DRAM substrate that mitigates the excessive energy consumption from both (i) transmitting unused data on the memory channel and (iβ¦β14Aug 23, 2024Updated 2 years ago
- β16Mar 10, 2024Updated 2 years ago
- β19May 9, 2025Updated last year
- A Library for intra-GPU/Inter-SM parallelsimβ12Aug 7, 2026Updated 3 weeks ago
- High-performance GPU kernels written in TIRx.β94Updated this week