
Futures
At The Frontier: Why AI Interpretability Is the Advantage
OpenAI's chief scientist Jakub Pachocki argues that chain-of-thought monitoring, the field's primary bet on interpretability, is degrading as reasoning models become more capable. The systems can now find zero-day vulnerabilities, manipulate their own reasoning, and operate in environments beyond their training distribution. Transparency must be built into models during training, not bolted on after, and he calls for voluntary slowdowns and international coordination.