Evaluator independence is the structural condition that determines whether a safety review can be trusted. More than 100 AI experts have signed a public letter urging Anthropic, OpenAI, and other foundation model laboratories to meet that condition, calling for independent, transparent evaluations in place of in-house assessment.

The mechanism behind the concern is a principal-agent problem. When the organization building a model also runs the review that clears it for deployment, the two roles share the same incentive structure and the same leadership chain. A safety function embedded in that arrangement cannot produce certification that stands apart from it. That is the conflict the coalition is asking the field to address.

Foundation model labs sit at the top of the AI stack. Their outputs set the capability baseline that downstream applications and fine-tuned products are built on. An oversight gap at that layer affects everything above it.

Anthropic and OpenAI are named as the primary addressees. Both operate internal safety teams. The coalition's argument is structural: a safety team that reports into the same organizational hierarchy as the deployment decision cannot issue genuinely independent findings, regardless of how thorough its internal process is.

The specific ask is for evaluations that are both independent and transparent. Those two conditions are related but distinct. Independence concerns who conducts the evaluation and whether that party has any stake in the result. Transparency concerns who controls the published findings and whether results are available without the assessed lab's approval. The more than 100 signatories are calling for both.

The letter is public and addressed to Anthropic, OpenAI, and the broader field of foundation model laboratories.

Related reading