NVIDIA/cutile-python

cuTile is a programming model for writing parallel kernels for NVIDIA GPUs

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 25 minutes ago
Added to GitGenius on December 17th, 2025
Created on June 13th, 2025
Open Issues & Pull Requests: 23 (+0)
Number of forks: 146
Total Stargazers: 2,133 (+0)
Total Subscribers: 21 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 33.7 hours
Mean response time: 7.8 days
90th percentile: 15.3 days
Tracked items: 69

How this project is maintained

Around half of the issues opened in the past year never receive a reply. Work labelled "status: needs-triage" is answered fastest, typically in under an hour, while "status: resolved" waits about 2 days. Only 7% of issues opened in the past year have been closed. Three people close 68% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 16
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 67 days
Stale 30+ days: 11
Stale 90+ days: 9

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 1
Events in 7 days: 1

Top labels

  • bug (29)
  • status: waiting-for-feedback (21)
  • feature request (19)
  • status: resolved (12)
  • status: triaged (12)
  • status: needs-triage (10)
  • dep: cuda-tileir (6)
  • documentation (3)

Detailed Description

cuTile Python is a programming language for writing parallel kernels on NVIDIA GPUs.

The tool addresses the challenge of GPU kernel development by providing a high-level programming model built on Tile IR, which abstracts away low-level GPU programming details. Developers write kernels in Python that are compiled to GPU code, allowing them to express parallel computation in terms of tiles—logical blocks of data—rather than managing individual threads directly. The approach requires NVIDIA Driver r580 or later and the tileiras compiler, which currently supports Blackwell and Ampere/Ada GPUs, with Hopper support planned for future versions.

The tool suits developers working on NVIDIA GPUs who want to write parallel kernels without dropping to lower-level languages. It is available through PyPI as the cuda-tile package for straightforward installation, or can be built from source for those needing the latest development version. Building from source requires a C++17 compiler, CMake, Python 3.10 or later, and CUDA Toolkit 13.1 or later. The project includes sample code and references TileGym for additional learning resources.

Maintainers typically respond to new issues and pull requests within a few days. Work in the issue tracker centers on bug reports, feature requests, and feedback collection.