safety-research/bloom

bloom - evaluate any behavior immediately  🌸🌱

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 17 minutes ago
Added to GitGenius on December 23rd, 2025
Created on June 24th, 2025
Open Issues & Pull Requests: 9 (+0)
Number of forks: 171
Total Stargazers: 1,390 (+0)
Total Subscribers: 13 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 25.7 hours
Mean response time: 2.7 days
90th percentile: 9.0 days
Tracked items: 17

How this project is maintained

Around half of the issues opened in the past year never receive a reply. Only 20% of issues opened in the past year have been closed. Three people close 100% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 0
New in 7 days: 0
Closed in 7 days: 0
Avg open age: N/A days
Stale 30+ days: 0
Stale 90+ days: 0

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Bloom is a Python tool for generating and running automated behavioral evaluations of large language models.

The tool addresses the challenge of systematically probing LLMs for specific behaviors like sycophancy, self-preservation, or political bias. Rather than relying on fixed benchmarks, Bloom takes a seed configuration describing a target behavior and generates diverse test scenarios tailored to that behavior. It then executes conversations with a target model and scores the results to measure behavior presence. The approach treats evaluation suites as reproducible artifacts tied to their seed configuration, allowing evaluations to be cited with full transparency about how they were constructed.

Bloom suits teams building safety evaluations or conducting interpretability research on language models. It works well for organizations that need to probe custom behaviors beyond standard benchmarks, or that want to test how stable a behavior is across variations—such as different noise levels or emotional pressure. The tool supports multiple model providers through LiteLLM integration and offers both conversation and simulated environment modalities. For large-scale experiments, it integrates with Weights and Biases for sweep-based evaluation across multiple models and configurations.

Development on this repository has been frozen at its last standalone release, with the project now maintained elsewhere. The tool remains usable as documented, but new projects should adopt the actively maintained version. The codebase provides pipeline stages that can be run individually—understanding, ideation, rollout, and judgment—allowing users to customize their evaluation workflow. An interactive viewer for results and an interactive chat mode for manual testing are included for exploration and debugging.