NewsNTech
Alignment sits at the base of every capability argument in frontier AI.
An Anthropic safety researcher has placed a specific probability on worst-case misalignment: a greater than 10% chance that AI could "kill all humans." The statement came after a colleague left the company over safety concerns.
The figure matters because Anthropic has positioned itself, more than most frontier labs, as a safety-first organization.
When a researcher inside that lab places a double-digit probability on catastrophic outcomes, it surfaces a tension that has followed every major capability release: the people building these systems are not uniformly confident they are safe.
Keep reading