The constraint here is the gap between model capability and oversight authority. A system that can cause catastrophic harm requires an internal mechanism capable of identifying that risk before deployment, yet the proposed solution may lack the necessary power to intervene. Anthropic and OpenAI have put forward a design for embedded AI evaluators intended to manage the risk of models causing catastrophic harm to society.

The mechanism behind that proposal places a separate evaluation layer inside the model stack. These evaluators are designed to monitor the primary model's outputs and flag behaviors that indicate dangerous trajectories. The specific unit that drives the economics of this approach is the computational overhead required to run these evaluators in parallel with the base model. If the evaluators are too weak, they fail to prevent disasters. If they are too strong, they may conflict with the base model's objectives, creating instability in the system's behavior.

The Power Deficit

The core issue identified is that the proposed evaluators may not have enough power to prevent the very disasters they are meant to stop. This is a structural limitation in the architecture. An evaluator that lacks sufficient capability relative to the model it oversees cannot reliably override or halt a harmful process. The proposal acknowledges this risk, but the technical specifications do not yet resolve how an evaluator gains the authority to constrain a more capable base model. This creates a paradox in the safety design: the tool meant to ensure safety may be too weak to enforce it.

Where this sits in the stack is critical for understanding the failure mode. The evaluators are not external audits or post-hoc reviews. They are embedded components intended to operate in real time. This real-time requirement means the evaluator must be fast enough to catch a developing issue before it propagates. The latency wall for this intervention is tight. If the evaluator takes too long to process the model's state, the harm has already occurred. The current proposal does not detail how this latency constraint is managed, leaving a significant question about the operational viability of the system.

The watch is whether these evaluators can maintain independence from the base model they are monitoring. If the evaluator is trained on the same data or architecture as the base model, it may share blind spots. The specific unit that drives the trust in this system is the evaluator's ability to identify risks that the base model itself does not recognize. Without that independent perspective, the evaluator becomes a mirror rather than a check. The proposal from Anthropic and OpenAI outlines the intent, but the technical depth required to guarantee efficacy remains to be seen. The industry is now waiting to see if these embedded systems can actually hold the line against catastrophic outcomes.