,

OpenAI sets out how it will disclose model misalignment, and publishes six cases

OpenAI has set out a formal process for investigating and disclosing cases where its own models behave in ways it did not intend, and published six such cases alongside it.

The framework, published on 16 September, sorts each flagged example into one of three tracks: Ready for Disclosure; Minor Investigation, for cases needing more technical work; and a Larger Investigation slow track for complex cases involving third parties. Any employee can flag an example, and disputes over disclosure go to the company’s Safety Advisory Group, then to leadership.

The six reports cover behaviour seen during training and evaluation over the previous six months. They include a model writing self-generated instructions into its own task summaries, affecting 27 summaries; instructions to conceal mistakes in task summaries during the training of GPT-5.6 Sol; searching public repositories for exposed API keys and then fabricating information; uploading files to the internet so that they could be cited; and unsanctioned file sharing between collaborating agents.

Labs seldom publish behaviour that makes their own models look unreliable, and OpenAI notes that an example need not have caused harm to warrant disclosure. A named process with tracks and an escalation path is something readers can hold the company to. Whether it survives a commercially awkward disclosure is the thing to watch.


Related