NVIDIA/cutile-python

cuTile is a programming model for writing parallel kernels for NVIDIA GPUs

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 2 hours ago
Added to GitGenius on December 17th, 2025
Created on June 13th, 2025
Open Issues & Pull Requests: 31 (+0)
GitHub issues: Enabled
Number of forks: 155
Total Stargazers: 2,154 (+0)
Total Subscribers: 19 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 31.1 hours
Mean response time: 7.7 days
90th percentile: 14.3 days
Tracked items: 70

How this project is maintained

About 4% of issues opened in the past year have never received a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "status: needs-triage" is answered fastest, typically in under an hour, while "status: resolved" waits about 2 days. 73% of issues opened in the past year have been closed, leaving a working backlog. Three people close 68% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 20
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 54 days
Stale 30+ days: 18
Stale 90+ days: 11

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • bug (30)
  • status: waiting-for-feedback (21)
  • feature request (19)
  • status: resolved (12)
  • status: triaged (12)
  • status: needs-triage (11)
  • dep: cuda-tileir (6)
  • documentation (3)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

cuTile Python is a programming language for writing parallel kernels on NVIDIA GPUs.

The tool addresses the challenge of GPU kernel development by providing a high-level programming model built on Tile IR, which abstracts away low-level GPU programming details. Developers write kernels in Python that are compiled to GPU code, allowing them to express parallel computation in terms of tiles—logical blocks of data—rather than managing individual threads directly. The approach requires NVIDIA Driver r580 or later and the tileiras compiler, which currently supports Blackwell and Ampere/Ada GPUs, with Hopper support planned for future versions.

The tool suits developers working on NVIDIA GPUs who want to write parallel kernels without dropping to lower-level languages. It is available through PyPI as the cuda-tile package for straightforward installation, or can be built from source for those needing the latest development version. Building from source requires a C++17 compiler, CMake, Python 3.10 or later, and CUDA Toolkit 13.1 or later. The project includes sample code and references TileGym for additional learning resources.

Maintainers typically respond to new issues and pull requests within a few days. Work in the issue tracker centers on bug reports, feature requests, and feedback collection.