Anthropic released new metrics detailing how much of its research is handled by its AI model, Claude. The company stated that Claude leads 26% of the work on future models, but the term 'lead' does not mean full autonomy.

The metrics show that as of August 2026, 26% of the work sits at AL4, up from under one percent in February. More than 90% reaches at least AL3. Claude hits AL5 nowhere.

Epoch AI calls AL4 'AI leads,' and Anthropic put that in its headline.

The company explained that at AL4, Claude can analyze, fix, and test a bug report without asking questions, but it isn't allowed to ship. A human reads the report and decides. The task and the direction still come from the human.

The difference from AL3 ('collaborates') is mainly that Claude no longer stalls when it runs into a snag.

Claude did the scoring itself. Agents gathered evidence from Slack and internal documents, and another Claude model assigned the levels. Anthropic admits this 'judge' could make the same mistakes as the system it's checking.

The share of tasks at level AL4 climbed from under one percent in February 2026 to 26 percent. Scored by Claude.

Where 'collaborates' ends and 'leads' begins isn't clear even to Anthropic. A cross-check by the company shows this: Employees were asked to judge how automated their own work area is.

When two people rated the same area, they landed on the same level only about a third of the time. The Claude model that assigned the official scores matched the human judgment 59 percent of the time.

Anthropic also didn't count work that advances safety and capabilities equally as safety. The company warns that the line to capability research is blurry, and every vendor is tempted to draw it generously. The burden of proof, it says, should sit with the developer.

Source: thedecoder