NVIDIA/cuda-rust

NVIDIA's CUDA platform for Rust. Host runtime crates plus Tile (cutile-rs) and SIMT (cuda-oxide) kernel programming models in idiomatic Rust.

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 1 hour ago
Added to GitGenius on September 28th, 2026
Created on April 22nd, 2026
Open Issues & Pull Requests: 152 (+0)
GitHub issues: Enabled
Number of forks: 302
Total Stargazers: 3,677 (+0)
Total Subscribers: 18 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 26.6 hours
Mean response time: 3.6 days
90th percentile: 10.3 days
Tracked items: 447

How this project is maintained

About 5% of issues opened in the past year have never received a reply. 95% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Almost all tracked open issues have seen activity in the last three months. 87% of issues opened in the past year have been closed, leaving a working backlog. Three people close 98% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 60
New in 7 days: 11
Closed in 7 days: 1
Avg open age: 30 days
Stale 30+ days: 9
Stale 90+ days: 0

Recent activity

Opened in 7 days: 9
Closed in 7 days: 1
Comments in 7 days: 4
Events in 7 days: 16

Top labels

  • bug (149)
  • codegen (141)
  • enhancement (72)
  • IR-lowering (60)
  • build-related (44)
  • intrinsics (34)
  • TBD (33)
  • cuda-feature (27)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

cuda-oxide is a Rust-to-CUDA compiler that lets you write GPU kernels in safe, idiomatic Rust and compile them directly to PTX without DSLs or foreign language bindings.

The tool solves the problem of writing CUDA kernels while staying within the Rust ecosystem and type system. It works by implementing a custom rustc backend that compiles functions marked with `#[kernel]` to CUDA PTX. The compilation pipeline converts Rust code through Rust MIR to a Pliron IR framework, then to LLVM IR, and finally to PTX. The project enables single-source compilation where host and device code live in the same file and build together with a single `cargo oxide build` command. Device-side abstractions include type-safe indexing, shared memory, scoped atomics, barriers, TMA operations, and warp and cluster operations. Kernel policies can be defined at compile time to create separate tuned specializations without runtime overhead. The host-side runtime handles memory management, pinned host transfers, and kernel launching through both synchronous and asynchronous interfaces.

Developers should adopt this tool if they want to write GPU kernels in pure Rust without learning CUDA C++ or managing separate compilation pipelines. It suits projects where the Rust ecosystem and type safety are priorities and where kernels can be expressed in idiomatic Rust. The project is in early alpha stage, so adopters should expect bugs, incomplete features, and API changes as development continues. The tool supports both checked and unchecked kernel launches, with the latter requiring unsafe code to prove correctness of launch dimensions and resources.

The project maintains active continuous integration for both the main codebase and example compilation. Development activity includes regular updates to the codebase and ongoing refinement of the API surface. The maintainers actively solicit community feedback and contributions to shape the project's direction.