NVIDIA-NeMo/Gym

Evaluate and improve models and agents using environments

View on GitHub ↗Jump to charts ↓

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 58 minutes ago
Added to GitGenius on December 17th, 2025
Created on August 25th, 2025
Open Issues & Pull Requests: 1,168 (+0)
GitHub issues: Enabled
Number of forks: 381
Total Stargazers: 1,227 (+0)
Total Subscribers: 8 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.2 days
Mean response time: 17.3 days
90th percentile: 53.2 days
Tracked items: 740

How this project is maintained

Roughly one issue in three opened in the past year never receives a reply. 97% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 37% of tracked open issues have had no activity in three months. Only 45% of issues opened in the past year have been closed. Three people close 53% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 601
New in 7 days: 33
Closed in 7 days: 14
Avg open age: 34 days
Stale 30+ days: 418
Stale 90+ days: 224

Recent activity

Opened in 7 days: 19
Closed in 7 days: 13
Comments in 7 days: 3
Events in 7 days: 54

Top labels

  • documentation (189)
  • feature (94)
  • community-request (80)
  • core-infra (79)
  • usability (48)
  • agents (46)
  • area:agent (40)
  • CLI (37)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

NeMo Gym is a library for evaluating and improving models and agents using environments. It provides infrastructure to develop environments, run evaluation and training at scale, and access a collection of popular benchmarks and training environments.

The tool addresses the need to evaluate models and agents in stateful, interactive settings where tasks require multiple steps and real-time feedback. An environment in NeMo Gym consists of a dataset defining tasks, an agent harness specifying how the model interacts with the world, a verifier that scores task completion, and state that tracks per-task execution context. This modular design allows developers to build reproducible evaluation systems that work consistently across teams and can scale to thousands of concurrent requests. The library handles the complexity of managing stateful interactions, making it suitable for scenarios like code execution, tool calling, and sandboxed environments where simple stateless scoring is insufficient.

Adopt NeMo Gym if you need reproducible evaluation across teams, require scale for multiple repeats per task or concurrent training requests, or want to seamlessly move between evaluation and agent optimization. The tool is less necessary if you only need to score model outputs with a stateless check and have no scaling requirements. NeMo Gym integrates with other environment libraries including Aviary, Harbor, OpenEnv, and Reasoning Gym, allowing you to combine benchmarks from multiple sources. It supports training with various RL frameworks and has been battle-tested in production Nemotron training.

The project maintains a substantial base of adopters who report issues from real-world use rather than the core team driving the issue tracker. Typical responses to issues or pull requests arrive within one to two weeks. Work in the issue tracker centers on documentation, core infrastructure, and community requests.