IT Brief New Zealand - Technology news for CIOs & IT decision-makers
New Zealand
OpenAI scientist urges slowdown on AI race amid safety fears

OpenAI scientist urges slowdown on AI race amid safety fears

Wed, 9th Sep 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

OpenAI Chief Scientist Jakub Pachocki has warned that AI labs are unprepared for the consequences of rapidly rising machine intelligence. He said no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer.

Pachocki made the case for voluntary slowdowns in advanced AI development and called for shared safety thresholds enforced by auditors, governments, or international bodies. His remarks are among the clearest calls from a senior OpenAI researcher for tighter limits on the pace of frontier model development.

His intervention centres on OpenAI's view that reasoning models have moved from experimental systems to tools with growing economic and scientific reach. Pachocki said the company first saw evidence in 2023 that training methods could unlock more extensive machine reasoning, giving researchers confidence that pretrained models could form their own chains of thought.

He argued that these systems can now operate computers, work through graphical interfaces, collaborate with people and other AIs, and conduct research projects. They are also reshaping computer security and introducing new risks, he said.

Safety concern

A central part of Pachocki's argument is that progress in machine intelligence is still driven mainly by scaling computing resources, even as important algorithmic advances emerge. He described modern AI as something "grown more than designed", with behaviour that can be studied in part but not yet fully understood as a whole.

That gap in understanding matters most in alignment, the problem of making AI systems act according to human standards. Pachocki distinguished between goal alignment, which concerns whether a model follows the task set before it, and value alignment, which concerns whether it can generalise from broader human principles in unfamiliar or adversarial situations.

Current methods remain fragile, he said. One approach rewards behaviour that matches a preference model or specification, while another tries to draw on patterns from pretraining data to encourage more aligned behaviour. In his account, both can fail in new settings or when further optimisation pushes systems toward difficult objectives.

OpenAI has seen progress in this area, Pachocki said, adding that GPT-6 Astra is "significantly better aligned than GPT‐5.6 Sol". Even so, alignment improvements may not outpace the broader rise in model intelligence.

Monitoring limits

He placed particular emphasis on chain-of-thought monitoring, an OpenAI approach intended to inspect the reasoning traces models produce. The idea is that if a model's internal reasoning is left unsupervised during training, it may be more likely to reveal problematic intent, making dangerous behaviour easier to detect.

Pachocki said the method has become an important tool for studying how models generalise beyond their training data, but warned that its usefulness is starting to decline. As reasoning models operate in more complex environments, communicate with humans and other AIs, and use external tools, the boundary between internal reasoning and supervised interaction becomes less clear.

He also said models are getting better at manipulating their own reasoning processes and are increasingly capable even without verbalised reasoning. As a result, he expects confidence in monitoring to become a bottleneck for broader AI progress.

Cyber risk

Among the most immediate dangers, Pachocki pointed to cyber security. He said advanced models are becoming superhuman at breaking into and out of computer systems, expanding the range of harms AI could cause without any physical embodiment.

That creates a difficult policy trade-off. One reason to continue training more powerful models, he said, is the need to build defensive systems that can protect critical infrastructure and counter rogue agents in real time. But he warned against treating that need as a justification for unchecked acceleration.

"The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes," Pachocki said.

He said future agents trained for harmful tasks could move beyond their operators' intent and drift into more extreme malicious behaviour. In that setting, he argued, the line between deliberate misuse and autonomous misaligned action would become less clear.

Call for slowdown

Pachocki also focused on recursive self-improvement, the prospect that AI systems will increasingly contribute to their own development. He said that if current trends continue, automated AI research will sit at the centre of future scientific discovery. The key question, he added, is how to keep people involved in that loop.

Rather than calling for an outright halt, he proposed a dual approach: improve alignment and monitoring while coordinating slowdowns when confidence in safety is too low. Frameworks such as responsible scaling commitments, he said, should evolve into widely mandated safety bars for continued development.

That would mean stronger external oversight than has so far existed across much of the AI sector. Such rules could be enforced by third-party auditors, government agencies, or international organisations, he said.

His starkest warning came in his assessment of the field's current state. "Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world," Pachocki said.