Sections

Commentary

In AI evaluation, access is not evidence

September 24, 2026


  • Key AI companies have just committed to giving outside evaluators a seat inside the building. But access and credible evaluation are not the same thing, and the hard part begins after evaluators get access.
  • A “measurement gap” is present: Tests are proliferating, but the field lacks the infrastructure that would let their results become comparable and cumulative evidence.
  • Establishing minimum reporting standards, comparing same system evaluations, building shared infrastructure, and validating evaluations against intended outcomes could turn greater access into better evidence.

Independent evaluators are getting inside AI companies. Now the field has to make their results mean something.

Two leading AI companies have just committed to giving outside evaluators a seat inside the building. On September 12, Anthropic’s Dario Amodei committed to giving independent evaluators “ongoing, employee-like access” to verify the company’s safety practices, with the right to publish without Anthropic’s editorial control. OpenAI’s Sam Altman said OpenAI would do the same. Independent researchers need that access. But access and credible evaluation are not the same thing, and treating the first as the second would waste this opportunity.

The hard part begins after evaluators get access: What do the resulting tests tell us? If two groups evaluate the same AI system and reach different conclusions, can we determine whether the difference came from the test, the evaluator, their level of access, ordinary variation, or a change in the system itself?

Too often, we cannot. Re-evaluations of widely used benchmarks have produced materially different scores and, in some cases, different model rankings. The same problem appears in frontier safety evaluation. In May 2025, Anthropic disclosed that Apollo Research had tested an early version of Anthropic’s Claude Opus 4 and advised against deploying it in situations where strategic deception is instrumentally useful, because of high rates of strategic deception. Anthropic reported that the behavior stemmed in part from a training data omission which Anthropic said it had since addressed release and that, based on its own audits, it believed the final model’s behavior in such scenarios was “roughly in line with other deployed models.” Anthropic publicly disclosed all of this. But because Apollo’s tests were not rerun on the final system, an outside reader could not determine how much of the difference reflected the reported correction rather than other differences between the evaluations. If the gap is visible even where disclosure is this extensive, it is unlikely to be narrower where it is not.

That is the measurement gap. Tests are proliferating—benchmarks, red teaming, safety and capability evaluations, audits, post-deployment monitoring—but the field lacks the infrastructure that would let their results become comparable and cumulative evidence.

When we developed the NIST AI Risk Management Framework, one of its four core functions was “Measure,” which asks for methods that are valid, reliable, and documented well enough that others can interpret the results. The field has adopted that vocabulary faster than the practice. The task now is to make AI measurement work more like measurement in mature fields, where results are interpretable, independently comparable, and validated against the outcomes they are meant to inform.

Independence, in particular, is more than distance. An external evaluator working under restrictive access or publication terms may have less practical independence than an internal team with technical autonomy and protected reporting rights. Useful questions are: Who controls the methodology? Who controls access? Who funds the work? Who controls publication? Those dependencies should be disclosed and, where possible, constrained. An evaluation is not made credible simply because someone outside the company ran it—or, for that matter, someone inside it.

4 steps to move from access to evidence

Four practical steps could turn greater access into better evidence. They apply to the whole evaluation ecosystem—academic labs, safety institutes, commercial auditors, developers’ own teams—not only to the evaluators now being embedded at two companies.

First, establish minimum reporting standards. Every evaluation report should state what was measured; the system and version tested, including configuration and safeguards; the protocol and conditions; the evaluator’s level of access; uncertainty and limitations; when the evaluation was conducted; and who conducted, funded, and controlled publication of the work.

Traceability—knowing what system was tested, how, and when—matters especially for AI, because models and the systems around them change continually through fine-tuning and updates to prompts, safeguards, tools, and retrieval. Funders, journals, standards bodies, and procurement offices can require these disclosures now, and government procurement is an especially powerful lever. This would standardize disclosure, not method. A young field should remain free to improve its methods while being required to say clearly what it did.

Second, compare evaluations of the same systems. Mature measurement fields routinely use interlaboratory comparisons: where multiple laboratories examine the same reference object so that disagreement reveals how much results depend on protocol, environment, instrument, or analyst. AI safety institutes, standards bodies, national measurement institutions, and other trusted intermediaries could convene analogous exercises. Different evaluators should test overlapping systems and questions, compare findings, and investigate disagreement. Where results converge across independent methods, confidence increases. Where they diverge, the differences become evidence about which choices matter and where methods need to improve.

The point is not simply to reproduce a score. Evaluators using the same benchmark, prompts, or automated grader can reproduce the same blind spots. Comparability also applies to qualitative work: surveys of red-teaming practice find wide variation in purpose, setting, and reporting. Recording the system, approach, coverage, effort, and findings would make it possible to tell whether teams observed the same kinds of behavior under comparable conditions. The goal is to make agreement and disagreement informative enough that one evaluation teaches us something about another.

Third, build shared evaluation infrastructure. Credible evaluation requires resources that individual evaluators cannot easily maintain and that developers should not control alone: these include held-out or refreshed test materials, secure environments for testing potentially dangerous capabilities, reference protocols that enable comparison over time, and common ways to record results. That infrastructure should test not only what a system can do but whether its safeguards work and how it behaves under realistic conditions. Increasingly, that means evaluating the system as deployed, not only the underlying model.

The hardest evaluations of frontier systems will still require developer cooperation and often developer compute. Developer cooperation may be necessary; developer control over methods, interpretation, or publication need not be. Some of these capabilities are public goods and will require public, philanthropic, or consortium funding structured so that no single developer can determine whether the infrastructure continues to exist or whether unfavorable findings are reported.

Fourth, validate evaluations against the outcomes they are intended to inform. Not every evaluation is meant to predict deployment outcomes; a capability test can be valid without predicting incidents. But when evaluations support claims about safety, reliability, or deployment risk, we should ask whether the evidence actually bears on those claims.

That means connecting pre-deployment evaluations with deployment experience. When an important failure occurs, did pre-deployment evaluation detect the relevant capability or failure mode? Underestimate it? Miss it entirely? Answering requires traceability after deployment as well as before, and enough information about the system, configuration, safeguards, and conditions to link what happened to what was tested.

This will be the slowest step. Incident reporting today is fragmentary, serious incidents may be rare, and when testing leads to mitigation we cannot directly observe what would have happened without it. Better incident reporting and system traceability are therefore prerequisites for validation. Even imperfect evidence can tell us where pre-deployment testing is informative—and where monitoring and other oversight must carry more of the load.

How we will know it is working

A year from now, take an important AI system and ask: Can independent groups produce results that can be meaningfully compared? Can a reader tell what each tested and under what conditions? Can important disagreements be explained? Where evaluations support deployment claims, does subsequent experience support those claims?

If so, AI evaluation will be starting to function as a measurement system.

The embedded evaluators now being placed at Anthropic and OpenAI are a natural place to start: They could be the first to work under common reporting standards, and their findings could be recorded in a form that outside groups can compare against. But they are one part of a much larger ecosystem, and the measurement gap runs through all of it. Access is an input, not the objective. The goal is not simply more evaluation. It is evidence we know how to interpret, compare, and trust.

The Brookings Institution is committed to quality, independence, and impact.
We are supported by a diverse array of funders. In line with our values and policies, each Brookings publication represents the sole views of its author(s).