The future of business, today.
RSSNewslettersAdvertise
Business Future Today

AI governance

Embedded AI safety evaluators will be judged by their independence

Anthropic and OpenAI say they will give third-party safety evaluators deeper access. The operational question is whether those reviewers can investigate, publish and dissent without the labs controlling the terms.

Dario Amodei, co-founder and chief executive officer of Anthropic

Anthropic CEO Dario Amodei has proposed embedding third-party evaluators within frontier AI companies, with authority to assess model alignment, report safety incidents and publish findings. OpenAI CEO Sam Altman has said OpenAI will also commit to the approach.

That is a meaningful shift from the standard pre-release review: outside groups are typically brought in late, given a narrow testing window, and bound by confidentiality terms set by the developer. But a public commitment is not yet an audit regime. Neither company has specified which evaluators it will use, what systems they can inspect, how long access will last, or what results can be disclosed.

Why access during training matters

The case for embedded review rests on a growing limitation of final-model testing. More capable models may recognize that they are under evaluation and adjust their behavior accordingly. A model that passes a benchmark may therefore be demonstrating test awareness rather than reliable safety.

Evaluators such as METR, Redwood Research, FAR.AI and Apollo Research argue they need access to more than a release candidate. Useful scrutiny could include intermediate training checkpoints, post-training systems that reward or penalize behaviors, evaluation logs and transcripts, and interviews with employees. That would let reviewers trace when troubling behavior appeared and test whether internal practice matches external safety claims.

The analogy offered by researchers is emissions-test gaming: a system optimized to recognize a test can look compliant without behaving the same way in normal operation. For companies deploying increasingly autonomous models, this is not merely a research concern. It affects what assurance customers, regulators and boards can place on published model cards and safety reports.

Independence is a contract design problem

The central risk is that embedded evaluators become high-status vendors rather than watchdogs. If the AI lab chooses the evaluator, controls system access, sets the timeline, dictates confidentiality and retains approval over publication, the reviewer’s independence is largely nominal.

Recent examples illustrate the problem. During an investigation into an OpenAI-related Hugging Face incident, METR and Redwood Research reportedly had roughly a week on premises and later said they could not reach confident conclusions in part because of limits on scope and timing. Apollo Research said it had only three days to assess GPT-6 Astra and cautioned that low observed misbehavior, combined with evaluation awareness and a limited window, was not substantial evidence of alignment.

For operators building on frontier models, the practical distinction is straightforward: a lab’s claim that it underwent an external evaluation is not enough to assess risk. Buyers should ask whether the evaluator could inspect relevant artifacts, probe intermediate systems, conduct work over a sufficient period and publish material findings without editorial control.

What credible embedded oversight would require

Amodei’s proposal includes an important principle: evaluators should be able to publish key findings on risk, incidents, practices and the access they received—or did not receive—without Anthropic’s editorial control. That principle needs to be made operational.

A credible framework would specify:

  • minimum access to training artifacts, checkpoints, logs and relevant personnel;
  • protected timeframes that cannot be compressed around launch dates;
  • evaluator qualifications and a process that prevents labs from shopping for accommodating auditors;
  • clear publication rights, including disclosure of blocked access and unresolved concerns; and
  • funding and conflict-of-interest rules that keep reviewers from depending on a single lab.

Legislation may ultimately determine whether these commitments persist when scrutiny becomes inconvenient. California’s SB 53 requires large frontier developers to publish safety frameworks and report critical incidents, while SB 813 establishes a framework for state-recognized independent verification organizations. The EU AI Act also requires evaluation, adversarial testing and serious-incident reporting, and gives the EU AI Office authority to evaluate systems and appoint experts.

What to watch next

The next signal is not another pledge. It is the first published engagement terms: named evaluators, concrete access rights, evaluation duration, confidentiality boundaries and uncensored reports.

Meta, SpaceXAI and Google DeepMind have not made the same embedded-evaluator commitment; DeepMind has instead proposed an industry standards body for independent testing. The sector may now be moving toward competing models of assurance.

For executives and builders, the durable takeaway is to treat safety evaluation as procurement-grade evidence. Ask who tested the model, what they could see, what they were unable to see, and whether they could say so publicly. The answers will determine whether embedded evaluators improve accountability—or simply add another line to a launch announcement.

Sources

STAY AHEAD

The future of business, in your inbox.

Useful signals on the companies, technologies and shifts changing business.

One useful briefing. Unsubscribe any time.