Claude is already building its own successors, but true autopilot remains a long way off
AI development is starting to become something of a circular process. Anthropic has revealed that Claude already “leads” 26% of the company’s AI research and development work. In February, the same figure was below 1%. More impressively still, AI is involved, at least as a collaborator, in more than 90% of Anthropic’s development work.
Before getting carried away with science fiction, however, it is worth defining the terms. Claude is not sitting in a digital lab and independently deciding which model to build next that will be smarter than itself.
Anthropic uses an automation scale from AL0 to AL5. AL4, or “leads”, means that AI receives a general task from a human and can complete most of it itself from start to finish. A human monitors the work but no longer has to specify every step in advance. AL5 would mean full autonomy without a human, and in Anthropic’s measured development work, Claude has not yet reached that level in any area.
The jump from 1% to 26% matters more than the catchy headline
The most notable thing is not even the current 26%, but the speed of the change. As recently as February, the share of work led by Claude was below 1%.
This creates a rather interesting feedback loop in AI development. A better model helps researchers and engineers build the next model, which in turn should be able to take on an even larger share of the work involved in developing the next generation.
Anthropic’s own recent example shows what such a division of labour means in practice. In less than four weeks, Claude optimised more than 30 open-source models for biomolecular modelling, making them around four times faster on average.
So this is no longer simply about a chatbot writing a few lines of Python code for an engineer. Agents can be given increasingly large, complete work packages.
About 30,000 AI agents work simultaneously in Anthropic’s lab
The scale makes the story even more interesting. On Anthropic’s most widely used internal platform, around 30,000 agents work simultaneously on research and engineering tasks. The company controls the agents’ activities using automated safety systems.
This is quite a good example of what “AI at work” is starting to look like in reality. One Claude chat window is not particularly exciting. Tens of thousands of agents solving tasks in parallel are.
Anthropic emphasises oversight because the larger the tasks given to agents, the more important the decisions they make during the work. At some point, the question is no longer whether the model wrote the right line of code, but which research direction it chose in the first place.
The figure should still be taken with a grain of salt
The 26% is not an independent audit or a universal industry standard. The metric was developed by Anthropic, and the company itself acknowledges that it uses its own models to assess the systems. This may create a situation in which the evaluating model makes the same errors as the model being evaluated. AI labs also currently lack a shared methodology that would allow Anthropic’s, OpenAI’s and Google DeepMind’s figures to be compared fairly side by side.
It would therefore be premature to declare that Claude is already doing a quarter of Anthropic researchers’ work or developing itself autonomously.
A much more accurate—and indeed more interesting—conclusion is this: in an increasingly large share of cutting-edge AI development, humans are moving from the role of hands-on practitioners to that of task-setters and supervisors.
If, according to Anthropic’s metric, Claude was able to lead less than 1% of development work in February and already 26% in August, this curve is the main thing to watch in future reports. AI that helps build the next AI is no longer a theoretical idea. For now, it simply still needs a human alongside it.