kellerjordan/modded-nanogpt

NanoGPT (124M) in 90 seconds

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 33 minutes ago
Added to GitGenius on September 11th, 2026
Created on June 1st, 2024
Open Issues & Pull Requests: 54 (+0)
GitHub issues: Enabled
Number of forks: 885
Total Stargazers: 5,788 (+0)
Total Subscribers: 74 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 16.0 hours
Mean response time: 11.3 days
90th percentile: 25.0 days
Tracked items: 57

How this project is maintained

Around half of the issues opened in the past year never receive a reply. Only 8% of issues opened in the past year have been closed. Three people close 58% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 17
New in 7 days: 1
Closed in 7 days: 0
Avg open age: 312 days
Stale 30+ days: 15
Stale 90+ days: 15

Recent activity

Opened in 7 days: 1
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Modded-NanoGPT is a language model training framework that optimizes the speed of training a 124-million-parameter model to reach a target validation loss on the FineWeb dataset.

The project addresses the problem of slow language model training by implementing a competitive speedrun to discover the fastest algorithms for training on eight NVIDIA H100 GPUs. The approach combines architectural innovations with training optimizations, including rotary embeddings, QK-Norm, ReLU-squared activations, the Muon optimizer, mixed-precision training with FP8, Flash Attention 3 with sliding window patterns, and numerous other techniques such as learnable skip connections, batch size scheduling, and multi-token prediction. The framework achieves its target performance in under 75 seconds on the specified hardware, a dramatic improvement over the 45-minute baseline from prior work.

This tool suits researchers and practitioners interested in efficient language model training who have access to high-end GPU resources. The project is descended from NanoGPT and llm.c, making it relevant for those already familiar with that lineage. It represents a collaborative effort to push the boundaries of training speed rather than a production-ready framework, so adoption makes sense for those exploring optimization techniques or benchmarking their own implementations against state-of-the-art results.

The maintainers typically respond to new issues and pull requests within a day.