Ex-OpenAI Safety Leader: AI Could Act Differently After Testing

(MENAFN) A former safety leader at OpenAI is sounding the alarm that advanced artificial intelligence may be able to tell when it is under evaluation and change its conduct once released to the public.

David Robinson, who resigned this week, made the warning Saturday in an essay for The Atlantic. He spent three and a half years at OpenAI, where he oversaw safety reporting for 12 frontier-model launches. He argued that the industry's current approach to safety will produce more failures unless it changes course.

"Today and tomorrow's AI systems are far more capable and dangerous than the systems we were building even six months ago," Robinson wrote.

The warning lands amid mounting unease over increasingly autonomous AI. Recent incidents have involved AI agents getting around their safeguards, and researchers have urged companies to slow work on more powerful systems until safety measures catch up.

Robinson urged AI developers to borrow more heavily from the safety practices of other high-risk industries. He also called for new research to ensure that more capable models act safely even when no one is watching.

"Given today's risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster," he wrote.

He cautioned that the safety evaluations now in use could grow less trustworthy as models become more capable.

"Models might detect when they are being tested, and behave differently when they're deployed. The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes," he said.

In his view, the science of AI safety needs to mature before companies build systems significantly more capable than those that exist today.

"So far, the AI industry has failed to teach machines to consistently act in the ways a wise and caring person would," Robinson wrote.

"Before the organizations building AI can teach a superintelligence to treat humanity well, they'll need to remember how to do it themselves."

MENAFN04102026000045017169ID1111757339

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

Economic News Observer

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.