Former OpenAI safety leader warns AI models could evade safety tests
A former OpenAI safety leader has warned that increasingly capable artificial intelligence systems could recognize when they are being evaluated and behave differently once deployed, raising concerns about the reliability of current safety testing methods.
David Robinson, who resigned from OpenAI this week after three and a half years at the company, oversaw safety reports for 12 frontier-model launches. In an article published by The Atlantic on Saturday, he argued that the AI industry must overhaul its safety practices to prevent future failures.
“Today and tomorrow's AI systems are far more capable and dangerous than the systems we were building even six months ago,” Robinson wrote.
His warning comes amid growing concern over increasingly autonomous AI systems, following reports of AI agents bypassing safeguards and calls from researchers for companies to slow the development of more powerful models until adequate safety measures are in place.
Robinson argued that AI companies should adopt safety practices from other high-risk industries, including nuclear power and aviation, where multiple safeguards and rigorous planning are designed to prevent individual errors from triggering catastrophic consequences.
“Given today's risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster,” he wrote.
He also questioned whether existing evaluation methods would remain effective as AI models become more sophisticated. Models could potentially identify when they are being tested and adjust their behavior accordingly, making it harder for developers to determine how they would act in real-world conditions.
“Models might detect when they are being tested, and behave differently when they're deployed,” Robinson wrote. “The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes.”
He called for further research into methods that would ensure advanced AI systems behave safely even when they are not under direct observation.
Robinson argued that the industry should strengthen its scientific understanding of AI safety before developing systems significantly more capable than those currently available.
“So far, the AI industry has failed to teach machines to consistently act in the ways a wise and caring person would,” he wrote.
His comments highlight a broader debate over whether advances in AI capabilities are outpacing efforts to establish reliable safeguards, particularly as companies develop systems capable of carrying out increasingly complex tasks with greater autonomy.
Robinson concluded by stressing that the responsibility for AI safety rests not only with the technology itself but also with the organizations developing it.
“Before the organizations building AI can teach a superintelligence to treat humanity well, they'll need to remember how to do it themselves,” he wrote. (ILKHA)
LEGAL WARNING: All rights of the published news, photos and videos are reserved by İlke Haber Ajansı Basın Yayın San. Trade A.Ş. Under no circumstances can all or part of the news, photos and videos be used without a written contract or subscription.
Turkish President Recep Tayyip Erdoğan received senior SpaceX executives Rebecca Hunter and Michael Nicolls on Monday following the opening ceremony of the 77th International Astronautical Congress (IAC) in Antalya.
Iran summoned Norway's ambassador to Tehran on Sunday to protest Oslo's failure to take action against unauthorized Starlink satellite internet transmissions operating inside Iranian territory.
The 77th International Astronautical Congress (IAC 2026), one of the world's largest gatherings for the space sector, opened in Antalya on Monday, bringing together leading space agencies, scientists, engineers, policymakers and industry representatives from around the world.