marin-community/marin

Open-source framework for the research and development of foundation models.

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 41 minutes ago
Added to GitGenius on August 26th, 2026
Created on March 22nd, 2024
Open Issues & Pull Requests: 567 (+0)
GitHub issues: Enabled
Number of forks: 248
Total Stargazers: 3,047 (+0)
Total Subscribers: 37 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 0.6 hours
Mean response time: 7.9 days
90th percentile: 12.8 days
Tracked items: 3,118

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 53% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. Almost all tracked open issues have seen activity in the last three months. Only 7% of issues opened in the past year have been closed. Three people close 59% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 474
New in 7 days: 42
Closed in 7 days: 65
Avg open age: 55 days
Stale 30+ days: 262
Stale 90+ days: 22

Recent activity

Opened in 7 days: 39
Closed in 7 days: 61
Comments in 7 days: 141
Events in 7 days: 324

Top labels

  • agent-generated (1,502)
  • stale (494)
  • bug (473)
  • p1 (468)
  • experiment (431)
  • p2 (328)
  • infrastructure (257)
  • tldr (204)

Detailed Description

Marin is an open-source framework for the research and development of foundation models.

Marin addresses the challenge of building large language models by providing infrastructure and methodology for every stage of the process: data curation, transformation, filtering, tokenization, pretraining, posttraining, and evaluation. The project's defining approach is radical transparency—it documents processes, experiments, and decisions as they happen, including failed attempts. This commitment to open development means that all process knowledge required to build foundation models is shared publicly, not just the final artifacts.

Marin suits researchers and organizations building large language models who value reproducibility and learning from the full development journey. The framework has proven flexible enough to support work beyond text, including audio-text models, DNA models, and protein models when used as a library. The project's current focus is pretraining a large mixture-of-experts model with over 500 billion parameters. For those interested in scaling laws and model recipes, Marin's Delphi scaling suite provides a complete reference implementation spanning from 3e18 to 1e23 FLOPs, including released checkpoints, training pipelines, recipe code, development methodology documentation, and plot-ready data.

Development activity shows sustained focus on frontier model training with active work on mixture-of-experts architectures and scaling research. The project maintains comprehensive documentation and makes intermediate artifacts available throughout the research process rather than only at completion. Contributions extend beyond the core team through library usage in external experiments, indicating the framework's utility as a foundation for specialized model variants.