cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
☆774Sep 7, 2026Updated this week
Alternatives and similar repositories for cutile-rs
Users that are interested in cutile-rs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Testbed for LLM inference with cutile-rs.☆71Jul 1, 2026Updated 2 months ago
- cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles stan…☆3,131Updated this week
- Multi-platform high-performance compute language extension for Rust.☆2,351Updated this week
- An Extensible Compiler IR Framework☆461Updated this week
- Safe rust wrapper around CUDA toolkit☆1,217Aug 12, 2026Updated 3 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-base…☆1,021Aug 28, 2026Updated last week
- 🐉 Making Rust a first-class language and ecosystem for GPU shaders 🚧☆3,331Aug 27, 2026Updated last week
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming☆807Updated this week
- Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.☆490Aug 17, 2026Updated 3 weeks ago
- Inference at the speed of light.☆2,969Updated this week
- Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.☆5,339Aug 17, 2026Updated 3 weeks ago
- Rust-native GPU kernel authoring framework: write GPU compute kernels in Rust, compile to PTX. The Triton equivalent for the Rust ecosyst…☆39Jun 12, 2026Updated 2 months ago
- cuTile is a programming model for writing parallel kernels for NVIDIA GPUs☆2,139Updated this week
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆60Feb 6, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- GPU based FFT written in Rust and CubeCL☆37Apr 24, 2026Updated 4 months ago
- The rustic MLIR bindings in Rust☆535Updated this week
- An Optimizer for Nvidia Compilers.☆132Aug 28, 2026Updated last week
- An attempt at safe imperative GPU programming.☆70Jul 6, 2026Updated 2 months ago
- CubeK: high-performance multi-platform kernels in CubeCL☆114Updated this week
- Type-level tensor shapes in Rust☆73Sep 1, 2026Updated last week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.☆932Updated this week
- Fast ML inference & training for ONNX models in Rust☆2,497Updated this week
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.☆15,881Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- 8-bit floating point types for Rust☆64Feb 4, 2026Updated 7 months ago
- A very fast linker for Linux☆3,952Updated this week
- Linear algebra foundation for the Rust programming language☆2,565Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,986Updated this week
- Tile primitives for speedy kernels☆3,667Aug 28, 2026Updated last week
- rust recursive deep learning framework☆61Aug 29, 2026Updated last week
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,482Updated this week
- A thread-per-core async Rust runtime with IOCP/io_uring/polling.☆1,870Updated this week
- Graph model execution API for Candle☆18Jul 27, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- PyTorch Single Controller☆1,076Updated this week
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.☆721Updated this week
- Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2☆676Updated this week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆381Jul 9, 2026Updated last month
- A Rust-native tensor & autodiff stack for scientific computing☆76Updated this week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 3 months ago
- Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.☆315Updated this week