Alignment sits at the base of every capability argument in frontier AI. An Anthropic safety researcher has placed a specific probability on worst-case misalignment: a greater than 10% chance that AI could "kill all humans." The statement came after a colleague left the company over safety concerns.

The figure matters because Anthropic has positioned itself, more than most frontier labs, as a safety-first organization. When a researcher inside that lab places a double-digit probability on catastrophic outcomes, it surfaces a tension that has followed every major capability release: the people building these systems are not uniformly confident they are safe.

The mechanism behind that concern is the alignment problem. Training a system to score well on human feedback does not guarantee the system will remain aligned with human values as its capabilities scales. The gap between benchmark performance and reliable safety under all conditions is where the risk lives.

The colleague's departure over safety concerns, alongside a sitting researcher's explicit greater-than-10% estimate, is a concrete data point about the internal state of one of the field's most prominent labs.

Related reading