Anthropic put permanent outside evaluators inside the company
Dario Amodei says third-party reviewers will get employee-level access as Anthropic calls for slower frontier development.
Verified 2:55 AM PDT · 3 original sources
Anthropic chief executive Dario Amodei published a plan Saturday to slow the rate at which frontier AI systems gain new capabilities. His first concrete step is a unilateral commitment to permanent third-party evaluators inside Anthropic. He says the reviewers should work for an outside organization, not Anthropic, and should receive company badges, desks and laptops. Their access would be mostly comparable to an internal risk-assessment team, with exceptions for legal or contractual limits.
Amodei names METR as an example of the kind of independent evaluator that could do the work. He does not announce a signed appointment. He also asks governments to require other frontier labs to provide comparable access. His broader proposal calls for common safety standards and limits on unchecked progress among leading companies in democratic countries. Government involvement would address antitrust concerns.
Amodei says two developments changed his view. He believes models have started contributing more directly to the next generation of AI development, raising the risk that capability gains outrun control work. He also points to the OpenAI-Hugging Face incident as evidence that current controls can fail outside the lab.
This is Anthropic's company proposal, not an enacted rule or an independent audit result. Amodei did not name a contracted evaluator, publish an agreement or identify a model release that Anthropic has delayed under the plan. Legal and contractual exceptions could limit access. Anthropic would still own the systems and make release decisions unless a future rule changes that authority.
Permanent access would move outside evaluation earlier in the development cycle. Most public model tests happen after a company chooses what to release and what evidence to publish. An embedded team could inspect systems while they are changing and compare internal and external risk judgments. It could document incidents before the public learns about them elsewhere.
The proposal also makes access a measurable promise. A call to slow down can remain a speech. Badges, laptops, system permissions and written reporting rights can be checked. Those details determine whether an evaluator can challenge a release decision or merely observe a process the company still controls.
Anthropic is asking competitors and governments to coordinate while still competing to build more capable systems. That conflict does not disappear because executives acknowledge it. A useful pacing rule needs a trigger, a decision maker, a duration and a public record of what happened when the trigger was reached.
Anthropic should name the evaluator, publish the access agreement and state when the embedded work begins. The agreement needs to define which models, training runs, incident logs and internal tests the team can inspect. It should also explain whether evaluators can publish findings without company approval and what happens when they recommend a delay.
Watch for a matching commitment from OpenAI, Google DeepMind, Meta or xAI. A shared standard also needs a government position on antitrust and enforcement. Without those pieces, one company's access policy will not set the pace for the rest of the market.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
Continue the edition