NVIDIA/cutlass

CUDA Templates and Python DSLs for High-Performance Linear Algebra

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 27 minutes ago
Added to GitGenius on July 16th, 2025
Created on November 30th, 2017
Open Issues & Pull Requests: 658 (+0)
Number of forks: 2,033
Total Stargazers: 10,291 (+0)
Total Subscribers: 126 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 16.5 hours
Mean response time: 25.9 days
90th percentile: 38.8 days
Tracked items: 1,074

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 39% of tracked open issues have had no activity in three months. Only 5% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 430
New in 7 days: 6
Closed in 7 days: 11
Avg open age: 360 days
Stale 30+ days: 378
Stale 90+ days: 335

Recent activity

Opened in 7 days: 2
Closed in 7 days: 9
Comments in 7 days: 10
Events in 7 days: 24

Top labels

  • ? - Needs Triage (1,127)
  • question (839)
  • inactive-30d (695)
  • inactive-90d (478)
  • bug (419)
  • CuTe DSL (205)
  • feature request (133)
  • CUTLASS C++ (76)

Detailed Description

CUTLASS is a collection of C++ template abstractions and Python domain-specific languages for implementing high-performance matrix multiplication and linear algebra computations on NVIDIA GPUs.

The tool addresses the challenge of writing efficient GEMM and related operations across diverse GPU architectures and data types. It decomposes these computations into reusable, modular components organized in a hierarchical parallelization structure. Developers can specialize and tune primitives at different levels through custom tiling sizes, data types, and algorithmic policies. The C++ template abstractions support an extensive range of data types including FP64, FP32, TF32, FP16, BF16, 8-bit floating point types, block-scaled formats, narrow integers, and binary types across Volta through Blackwell architectures. The newer Python DSL interface, built on core CUTLASS and CuTe concepts, enables high-performance kernel development without C++ expertise, offering faster compile times and native integration with deep learning frameworks.

CUTLASS suits performance engineers, researchers, and students working on GPU-accelerated linear algebra. It is particularly valuable for projects requiring specialized data types, mixed-precision computations, or fine-grained control over tensor core utilization. The C++ abstractions appeal to those needing maximum flexibility and control, while the Python DSL provides a lower barrier to entry for rapid prototyping and kernel design without sacrificing performance.

The project receives issue reports predominantly from external users rather than the core team, reflecting a substantial adopter base. Maintainers typically respond to new issues and pull requests within a day. Work in the issue tracker centers on triage, user questions, and inactive items, indicating active engagement with incoming user feedback.