Anthropic says Claude now leads 26% of its AI research
The company reports a sharp rise in research led by Claude. Its definition still includes human supervision.
Verified 12:21 AM PDT · 2 original sources
Anthropic released a new way to measure research automation on September 17. By its count, Claude led 26% of its AI research and development work in August, up from less than 1% in February. For Anthropic, leading means doing most of a task from a broad instruction while a person supervises.
The company says none of the work it measured was fully autonomous. More than 90% involved at least substantial collaboration with AI. Anthropic measured its own work. It did not independently test how quickly AI improves itself.
These are company measurements, not an independent audit. Anthropic uses its own models in evaluating its systems. It acknowledges that a model judging another model can share the same errors. Different labs also lack a common measurement method.
When AI does most of a research task, people must review work they did not do themselves. A count of finished tasks cannot tell them whether the results are right.
Anthropic's index could show how much work Claude takes on. If that share rises, the index still cannot tell whether the research gets better, mistakes increase, or reviewers keep up.
Anthropic plans to let outside evaluators see its internal systems and data. Their findings could make the index more useful if Anthropic measures future work the same way.
Those evaluators still need to check the ratings and the research behind them.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
Continue the edition