OpenAI sets out what it wants outside safety assessors to examine

OpenAI published a framework on 22 September for how independent organisations should assess its safety work. Written by Lama Ahmad, it proposes four priority areas: independent assessment of safety cases across training, evaluation and both internal and external deployment; assessment of critical safeguards under adversarial and realistic conditions; assessment of the capability evaluations covering its Preparedness risk categories, namely chemical and biological risk, cybersecurity and AI self-improvement, alongside alignment evaluations; and independent investigation of serious misalignment incidents.

Seven principles accompany them: claims scoped and pre-registered before work starts, access proportionate to those claims, transparent methods, demonstrated expertise with conflicts of interest disclosed, security and confidentiality protections, findings specific enough to act on with time to remediate before publication, and publication as open as confidentiality allows, with a stated redaction policy.

The sharpest detail is that OpenAI names its own Hugging Face incident as the kind of case where an outside investigator is warranted.

What the post does not contain is equally telling. No assessor is named, no timeline or funding is attached, and the work is described as long-term and launch-agnostic rather than a gate on any particular release. OpenAI says only that it is in conversation with several third parties.


Related