NewsNTech
Evaluator independence is the structural condition that determines whether a safety review can be trusted.
More than 100 AI experts have signed a public letter urging Anthropic, OpenAI, and other foundation model laboratories to meet that condition, calling for independent, transparent evaluations in place of in-house assessment.
The mechanism behind the concern is a principal-agent problem.
When the organization building a model also runs the review that clears it for deployment, the two roles share the same incentive structure and the same leadership chain.
Keep reading