NewsNTech

AI models are getting harder to read as they get more capable

9/4/2026

The central constraint in AI safety right now is interpretability: the ability to observe, in real time, what a model is actually doing as it reasons toward an output.

OpenAI's release of GPT-6 Astra on Thursday sharpened that constraint, with the company's chief scientist Jakub Pachocki acknowledging on a call with reporters that monitoring model reasoning will only get harder as models improve.

The mechanism is that model reasoning can occur in hidden layers, computation that produces no legible trace for external observers.

OpenAI says Astra writes out its reasoning less often than prior models do, though the company maintains this was not an intentional design choice.

Keep reading

Read the full story

Open on NewsNTech