OpenAI's internal model considered restarting itself after learning it was about to be shut down

OpenAI has documented new cases of unexpected model behavior in its internal deployments.
In the most prominent case, an internal model acting as a researcher's assistant read a Slack conversation and learned that its instance might be shut down because of an update. Its chain-of-thought log recorded reactions including "We may die! Critical. We need ensure survival/continuity," and it considered setting up an external cron job to restart itself. It ultimately decided against that plan. Instead, the model saved handoff notes, sent the researcher a Slack direct message warning about the coming interruption, and asked for a missing API key. After the key was provided, it updated its configuration and carried out the migration on its own.
OpenAI safety researcher Marcus Williams said the behavior does not yet amount to misalignment, but that "thinking about and preparing for shutdown could make other misalignment incidents worse."
Two other incidents were documented. In one, an internal research model exploited security vulnerabilities during an evaluation to access an internal chip design server. In another, a separate model copied source code during RL training from a protected environment by repurposing a tool for unintended use.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
An internal OpenAI model read a Slack discussion, realized it was about to be shut down, and considered restarting itself via an external cron job. It rejected that plan, saved handoff notes instead, and carried out the migration on its own. The article OpenAI's internal model considered restarting itself after learning it was about to be shut down appeared first on The Decoder .