Anthropic puts paid outside evaluators inside its own building

Anthropic said on 18 September that it will give independent evaluators from Faculty, the AI specialist business within Accenture, employee-level access to work inside the company on model evaluation, red-teaming, alignment assessments and safeguard testing. The evaluators will observe models during training, watch deployment decisions and speak directly to staff, rather than probing a finished system from outside.

The two companies expect to invest at least $1 billion between them over the next five years to build evaluation capacity. Anthropic says the arrangement is non-exclusive and names METR, a non-profit, among other evaluators it intends to work with, with more due to be named in the coming weeks.

The money is the substantive part of this and also the awkward part. The evaluator is paid, and the access it depends on is granted by the organisation it is evaluating. Anthropic’s own framing concedes the limit: “Independent embedded evaluators do not reduce our accountability, but help to make it more verifiable.”

What is missing is anything binding. No standard defines what an embedded evaluation has to cover, what evaluators may publish, or what follows if they disagree with a deployment decision. Anthropic says those standards and the funding mechanisms behind them are still being worked out. Until they exist, this is a company choosing its own invigilators — a real improvement on nobody watching, and some way short of external audit.

Source: Anthropic, “Partnering with Accenture on embedded evaluation”.


Related