Translating CUDA Tile Operations from Python to Rust Using Agentic AI
NVIDIA has published an account of a multi-agent workflow that translates tile-based GPU kernels written in cuTile Python and Triton-TileIR into cuTile Rust. The skill, named tilegym-converting-python-to-rust, ships in the NVIDIA/TileGym GitHub repository.
Using the skill, the team reports porting all 24 public TileGym operators to cuTile Rust, comprising roughly 40 GPU kernels, from element-wise operations to flash-attention decode, Multi-head Latent Attention (MLA) and mixture-of-experts (MoE) models. Some operators require multiple kernel variants. Average performance reached 99.5% of cuTile Python, and NVIDIA says all 24 operators clear a 0.95 geomean speedup threshold versus cuTile Python when benchmarked on NVIDIA DGX B200.
cuTile Rust, also called cutile-rs, is described as a tile-based system for safe, idiomatic GPU kernel authoring in Rust. It extends the Rust ownership model to tile-based GPU kernels by splitting mutable outputs into disjoint pieces and preserving the host-side ownership contract across kernel launches. Programmers can opt out locally for lower-level control, enabling direct execution of Tile IR operations.
Converting a kernel starts from whichever reference implementation an operator has and passes through a bounded multi-agent pipeline covering analysis, the device kernel, host and FFI code, and benchmarking. Each stage ends in a machine-checkable verdict, with validator scripts and Tile IR diffs deciding whether a conversion moves forward. Because cuTile Python, Triton-TileIR and cuTile Rust all emit the same CUDA Tile IR dialect and feed the same tileiras compiler, NVIDIA says a faithful port can be checked structurally by dumping and diffing reference and translated IR before tests run.
A central translation challenge is that cuTile Python JIT specializes each kernel implicitly at call time, while Rust requires every specialization to be declared in the kernel signature. Converted kernels integrate with TileGym through a C-ABI layer that passes tensor descriptors without copying or allocating, and the launcher never takes ownership of tensors.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
cuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to...