cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
☆742Aug 18, 2026Updated this week
Alternatives and similar repositories for cutile-rs
Users that are interested in cutile-rs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Testbed for LLM inference with cutile-rs.☆72Jul 1, 2026Updated last month
- cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles stan…☆3,080Updated this week
- Multi-platform high-performance compute language extension for Rust.☆2,321Updated this week
- An Extensible Compiler IR Framework☆452Updated this week
- Safe rust wrapper around CUDA toolkit☆1,207Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-base…☆1,008Jul 22, 2026Updated 3 weeks ago
- 🐉 Making Rust a first-class language and ecosystem for GPU shaders 🚧☆3,293Updated this week
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming☆797Updated this week
- Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.☆491Updated this week
- Inference at the speed of light.☆2,935Updated this week
- Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.☆5,320Updated this week
- cuTile is a programming model for writing parallel kernels for NVIDIA GPUs☆2,128Updated this week
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆61Feb 6, 2026Updated 6 months ago
- GPU based FFT written in Rust and CubeCL☆36Apr 24, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Rust-native GPU kernel authoring framework: write GPU compute kernels in Rust, compile to PTX. The Triton equivalent for the Rust ecosyst…☆37Jun 12, 2026Updated 2 months ago
- The rustic MLIR bindings in Rust☆533Aug 7, 2026Updated last week
- An Optimizer for Nvidia Compilers.☆128Updated this week
- An attempt at safe imperative GPU programming.☆70Jul 6, 2026Updated last month
- CubeK: high-performance multi-platform kernels in CubeCL☆111Updated this week
- Type-level tensor shapes in Rust☆73Jul 31, 2026Updated 2 weeks ago
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.☆922Updated this week
- Fast ML inference & training for ONNX models in Rust☆2,456Updated this week
- 8-bit floating point types for Rust☆64Feb 4, 2026Updated 6 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.☆15,765Updated this week
- Linear algebra foundation for the Rust programming language☆2,564Jun 24, 2026Updated last month
- A very fast linker for Linux☆3,857Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,795Updated this week
- rust recursive deep learning framework☆60Updated this week
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,431Aug 4, 2026Updated 2 weeks ago
- Tile primitives for speedy kernels☆3,633Jul 13, 2026Updated last month
- A thread-per-core async Rust runtime with IOCP/io_uring/polling.☆1,844Updated this week
- Graph model execution API for Candle☆18Jul 27, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- PyTorch Single Controller☆1,073Updated this week
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.☆714Updated this week
- Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2☆650Updated this week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆374Jul 9, 2026Updated last month
- A Rust-native tensor & autodiff stack for scientific computing☆70Updated this week
- Early-stage Rust drop-in alternative frontend for vLLM☆73May 22, 2026Updated 2 months ago
- Hybrid in-memory and disk cache in Rust☆1,788Updated this week