Bloom turns a hypothesis about AI behavior into a reproducible evaluation
Anthropic has opened a system that generates scenarios, runs conversations, and judges whether a model displays a defined behavior. Its value depends on preserving the seed, checking the judge, and not mistaking elicited behavior for real-world incidence.