pytorch/torchtitan

A PyTorch native platform for training generative AI models

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 2 hours ago
Added to GitGenius on May 13th, 2025
Created on December 13th, 2023
Open Issues & Pull Requests: 227 (+1)
GitHub issues: Enabled
Number of forks: 1,025
Total Stargazers: 5,791 (+0)
Total Subscribers: 59 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 4.5 hours
Mean response time: 19.7 days
90th percentile: 22.6 days
Tracked items: 782

Maintainer activity

41 people did triage or write work on this repository in the last 12 months.

At least 24% of torchtitan's 41 maintainers work at Meta. 23 say where they work, and 10 of those are Meta.

PyTorch Foundation owns this repository. 3 of its people did this work here, and are not included above.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 4% of issues opened in the past year have never received a reply. 95% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "module: checkpoint" is answered fastest, typically in about 3 hours, while "release blocking" waits about 6 days. 42% of tracked open issues have had no activity in three months. 76% of issues opened in the past year have been closed, leaving a working backlog.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 85
New in 7 days: 8
Closed in 7 days: 158
Avg open age: 22 days
Stale 30+ days: 60
Stale 90+ days: 37

Recent activity

Opened in 7 days: 6
Closed in 7 days: 158
Comments in 7 days: 6
Events in 7 days: 178

Top labels

  • question (105)
  • bug (71)
  • enhancement (70)
  • triage review (66)
  • high priority (61)
  • module: torch.compile (29)
  • documentation (19)
  • module: checkpoint (18)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

TorchTitan is a PyTorch native platform designed for rapid experimentation and large-scale training of generative AI models. It serves as a minimal clean-room implementation of PyTorch native scaling techniques, providing a flexible foundation for developers to build upon. The project is currently under extensive development, with the README recommending users employ the most recent PyTorch nightly to access the latest features. The platform has achieved significant milestones, including acceptance of its research paper at ICLR 2025, a GPU MODE lecture in December 2024, and a presentation at PyTorch Conference 2024.

The core mission of TorchTitan is to accelerate innovation in generative AI by empowering researchers and developers to explore new modeling architectures and infrastructure techniques. The platform is designed around three guiding principles: ease of understanding, use, and extension for different training purposes; minimal changes to model code when applying multi-dimensional parallelism; and a bias towards a clean, minimal codebase while providing basic reusable and swappable components. TorchTitan has been showcasing PyTorch's latest distributed training features through support for pretraining Llama 3.1 LLMs of various sizes, including 8B, 70B, and 405B parameter models.

The platform offers comprehensive support for multi-dimensional composable parallelisms, including FSDP2 with per-parameter sharding, Tensor Parallel with async variants, Pipeline Parallel with zero-bubble capabilities, and Context Parallel for training long-context LLMs. Additional key features include meta device initialization, selective and full activation checkpointing, distributed checkpointing with async support, torch.compile integration, Float8 and MXFP8 quantization support, supervised fine-tuning with chat-formatted datasets, and flexible learning rate scheduling. The platform provides extensive debugging tools, structured logging, and helper scripts for tokenizer downloads, checkpoint conversion, and distributed inference.

TorchTitan supports installation through multiple methods including direct source code execution, nightly builds, and stable releases via pip or conda. The platform includes an experiments folder to accelerate contributions and innovations, with dedicated guidelines for both experimental contributions and core fixes. Performance has been reported on up to 512 GPUs, with verified loss convergence correctness across various techniques. The source code is made available under a BSD 3 license, though users may have other legal obligations governing their use of linked third-party data and models.