β21Mar 17, 2026Updated 4 months ago
Alternatives and similar repositories for redfuser
Users that are interested in redfuser are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β24Sep 1, 2025Updated 11 months ago
- πAutomatically Update LLM inference systems Papers Daily using Github Actions (Update Every 12th hours)β12Updated this week
- A DL compiler fuzzerβ15Nov 1, 2024Updated last year
- A Triton-only attention backend for vLLMβ28Jul 14, 2026Updated 3 weeks ago
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitiveβ75Updated this week
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- β19Mar 4, 2025Updated last year
- β16Jan 24, 2024Updated 2 years ago
- DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scalingβ35Updated this week
- β27Feb 20, 2024Updated 2 years ago
- OSDI 2023 Welder, deeplearning compilerβ35Nov 24, 2023Updated 2 years ago
- Autonomous CUDA kernel optimization agent with structured task specs and per-config scoringβ17Jun 17, 2026Updated last month
- β24Jun 10, 2026Updated 2 months ago
- Verified graph rewriting (for dataflow circuits).β27Jul 15, 2026Updated 3 weeks ago
- β30Apr 7, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Multiple 1-stencil implementations using nvidia cuda.β12Dec 2, 2017Updated 8 years ago
- Simple python library for generating your own perfetto traces for your application. Can be used for both app instrumentation and custom β¦β27Jun 22, 2025Updated last year
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batchβ¦β129Jun 29, 2026Updated last month
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.β19Jun 12, 2026Updated last month
- β27Oct 1, 2025Updated 10 months ago
- Discovery of Structured Parallelism In Sequential and Parallel Codeβ10Feb 13, 2021Updated 5 years ago
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Codeβ16Mar 19, 2023Updated 3 years ago
- β33Jul 17, 2024Updated 2 years ago
- β10Mar 8, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- β154Updated this week
- GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weavingβ21Jul 30, 2025Updated last year
- Public benchmark results from Kernel Arena, a leaderboard for LLM-generated AI accelerator kernels.β21Mar 11, 2026Updated 4 months ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and minβ¦β10Aug 13, 2024Updated last year
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.β117Updated this week
- β194Updated this week
- Website for CSE 234, Winter 2025β16Mar 24, 2025Updated last year
- SIMPLER MAGIC: Synthesis and In-memory MaPping of Logic Execution in a single Row for Memristor Aided loGICβ13Dec 5, 2019Updated 6 years ago
- MIT 6.172 Performance Engineering of Software Systemsβ16Dec 30, 2021Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Share your GPU without MIG or MPSβ49Jan 27, 2026Updated 6 months ago
- β46Oct 15, 2025Updated 9 months ago
- A collection of out-of-tree extensions for the Triton language and compilerβ33Updated this week
- A new DRAM substrate that mitigates the excessive energy consumption from both (i) transmitting unused data on the memory channel and (iβ¦β14Aug 23, 2024Updated last year
- β16Mar 10, 2024Updated 2 years ago
- β19May 9, 2025Updated last year
- A Library for intra-GPU/Inter-SM parallelsimβ12Updated this week