NewsNTech
Recursive self-improvement is the mechanism that AI researchers at Anthropic and OpenAI are now describing as existential: the process by which an AI system extending its own capabilities becomes progressively better at further self-improvement, with the speed of that cycle determining how much room humans have to intervene.
Researchers at both organizations are warning that as that cycle accelerates, advanced systems could become substantially harder for humans to control.
A slow improvement cycle stays within the bandwidth of human evaluation. Engineers and safety researchers can observe a change and evaluate it before the system changes again.
When the cycle accelerates, that evaluation window contracts with each pass. The control mechanisms designed for one rate of improvement may not hold at a faster one.
Keep reading