Anthropic has proposed three measurements it says frontier AI labs should publish regularly, and released its own figures against them.
The first tracks how much of a lab’s AI research is done by AI, on a scale from AL0 (no AI involvement) to AL5 (fully autonomous). As of August 2026, Anthropic reports Claude “leads” 26% of its AI research and development work — completing most of a task end to end from a high-level prompt — up from under 1% in February 2026. More than 90% sits at or above the weaker “AI collaborates” level.
The second measurement is agent oversight. Anthropic says 0.002% of agent decisions were blocked by its monitor, roughly one in 47,000. The third is compute allocation: about 6% of the compute that went to AI R&D was spent on safety work, rising to about 12% for AI-driven AI R&D.
The caveats are the interesting part, and Anthropic states them itself. Safety research is hard to distinguish from capabilities research, it notes, and every developer “will be tempted to draw the line generously”. Its internal usage labels are best-effort and unverified. And it is using its own models to judge its own systems, so the judge may repeat the errors it is checking. It asks rival labs to publish comparable figures, independently verified.
