CubeCL

One codebase.More hardware choices.

Write your compute kernels in Rust. CubeCL translates them for different GPUs and CPUs, so you can reuse the same logic across hardware backends.

Open source · Apache 2.0 / MIT · By Tracel

Explore six alternative hardware paths through CubeCL. Open Rust to read the illustrative square.rs kernel using the 0.10 API. Equal-length buffers, device setup, and launch code are required to run it. The illustration does not execute or compile Rust.

One Rust kernel on CubeCL connects to NVIDIA via CUDA, AMD via HIP, Apple via Metal, Vulkan, WebGPU and CPU. Each path is an alternative hardware target.
  1. 01Write in Rust
  2. 02Specialize & compile
  3. 03Run on your device

Core features

Get more from your hardware.

01 / Rust

Write kernels in real Rust.

Use traits, generics, structs, and enums inside kernels to compose reusable building blocks at no runtime cost.

Explore language support

02 / Comptime

Compile for the real workload.

Kernels compile on first launch for the real shapes and device, and use specialized hardware such as tensor cores when it's available.

Explore comptime

03 / Autotuning

Tune on your own device.

CubeCL benchmarks candidate kernels on your hardware, keeps the fastest, and caches that choice for future runs.

Explore autotuning

Get started

All third-party product names and logos are trademarks or registered trademarks of their respective owners. They are used for identification only and do not imply affiliation with or endorsement of Tracel. WebGPU logo by W3C, used under CC BY 4.0.