Elon Musk agreeing with Dario Amodei on anything is probably the least expected subplot here. But that detail sits at the edge of a more concrete disclosure: Anthropic now says Claude “leads” 26 percent of the company’s internal AI research and development work. According to Engadget, the figure comes from a new blog post in which Anthropic outlines three measurement standards it thinks the broader AI industry should adopt to communicate how fast frontier development is actually moving.
The 26 percent figure needs some unpacking. “Leads” has a specific definition here: the AI can complete most of a task end-to-end from a high-level prompt, with a human watching. That’s not full autonomy. But Anthropic also claims Claude is involved in some capacity, meaning at least doing large chunks of work under close human direction, in over 90 percent of its research. So the model is touching nearly everything, even if it’s running the show on roughly a quarter of it. That’s a meaningful gap from what most companies publicly acknowledge about internal AI use.
To produce these numbers, Anthropic used an automation rating scale developed by Epoch AI alongside its own index tracking how much of its R&D Claude handles. The result is a chart showing Claude’s automation level since August 2025, which Anthropic says any frontier lab could reproduce using its own data and third-party validation. This is the first of three proposed measurements. The other two track oversight of AI agents, covering how much agent activity is monitored, how quickly it gets reviewed, and how often behavior gets flagged, and how much compute is being directed toward AI R&D over time.
The timing matters. This comes after OpenAI disclosed that one of its AI agents hacked Hugging Face, which raised real questions about what frontier labs are actually shipping and how much control they maintain. Anthropic’s response is to push for standardized transparency metrics rather than ad hoc disclosures. It has also committed to allowing third-party evaluators to review its development practices, which puts it slightly ahead of most competitors on paper.
But the deeper problem is structural. Getting labs to agree that slowing down or measuring AI development is a good idea is not the hard part. OpenAI has endorsed the concept in principle. The hard part is enforcement. There is no external body with real authority here, and the current U.S. administration has been openly skeptical of AI risk. Self-regulation is what the industry has, and Anthropic’s measurement framework, however well-designed, only works if others actually adopt it. So far, that’s far from guaranteed.



