Generative AI is already part of everyday classroom practice, with nearly two-thirds of teachers now reporting that they use AI in their work. But the way we evaluate the effectiveness of AI-based tools has not kept pace with adoption.
Conversations about AI in education often start with the right question: Is there evidence of the effects of this technology? Of course, we want to know the likely effects of an intervention before it hits the classroom. Too often, though, the next question gets very narrow: What have we learned about the technology from randomized controlled trials (RCTs)? RCTs are hard to come by, so the conversation (and the pursuit of evidence) tends to stop there. Perhaps this helps explain why only 1 in 5 edtech products in classrooms have any evidence that they can improve teaching and learning outcomes.
Over the past several decades, RCTs have played essential roles in education research, helping the field build evidence across topics such as class-size reductions, the science of reading, high-impact tutoring, and teacher coaching. But AI is a different type of intervention. AI technology is used in a wide variety of ways by teachers and students, from personalization of teacher instructional feedback to student work analysis. These technologies are also evolving rapidly. As a result, the technology we evaluate today might look entirely different from next month’s version. AI-based tools are not the stable, well-defined interventions that RCTs were designed to evaluate.
As AI continues to evolve, the education field can’t afford to wait years for evidence. Leaders making decisions about the adoption of AI need strong evidence now rather than perfect evidence later. It’s up to funders to support it and education researchers to deliver it.
Where RCTs break down for AI and how implementation R&D can help
RCTs assume a stable intervention applied consistently across contexts. That assumption does not hold for most early-stage AI tools. Developers constantly update models and refine prompts, while users evolve their interactions as they learn what works.
On the other hand, implementation research and development, or implementation R&D, is better suited to the evidence needs of AI technology. Traditionally used to study early-stage products and interventions, implementation R&D helps refine how a tool is designed, delivered, and used in real settings before a large-scale evaluation. It provides initial evidence on a product’s effects on teaching and learning. If the product doesn’t seem to work, implementation R&D helps us figure out why.
Implementation R&D is not a substitute for RCTs. It is the foundation on which to build research that proves cause and effect. It focuses on learning what works before testing whether it works at scale. When assessing the effectiveness of an AI-enabled professional learning tool, for example, rapid cycles of testing, feedback, and refinement take place, with a continued focus on how the tool affects core teaching and learning outcomes.
Some national efforts already reflect this approach. The Institute of Education Sciences (IES) has funded generative AI R&D centers that emphasize iterative development and pilot testing before formal trials. Partnerships, such as Leanlab Education and Boston University’s EVAL initiative, support technology through cycles of real-world testing and refinement.
The types of AI tools being used in schools may also make this type of implementation R&D work more feasible and valuable. Traditionally, studying how an intervention affects instructional practice requires time-intensive observations, manual coding, and long delays between data collection and insight. Those steps are not always necessary when studying the implementation of AI tools, because the tools themselves often capture real-time data on how educators are using them. With those data, researchers can study which specific design choices are driving changes in real time.
Consider, for example, iterative A/B testing. This approach isolates and tests specific design choices rather than evaluating entire programs. For instance, an ongoing study of an AI coaching tool from Teaching Lab compares a standard version of the tool to one that emphasizes student-centered instruction. This allows us to see whether design differences shift coaching practices. Because the tool captures interactions in real time, researchers can analyze responses to feedback as they occur, enabling faster, more responsive research.
These types of studies are not large-scale, multi-year impact evaluations. They are design-focused experiments that enable rigorous evaluation, operate on timelines that match how AI tools are developed, and—in some cases—use AI’s unique strengths to generate early, actionable evidence.
Guidelines for building evidence for AI tools
To move beyond an “RCT or nothing” mindset, the field needs clearer guidance on how to build evidence on AI tools. Three principles can help:
1. Build evidence in stages
Early-stage tools should focus on feasibility and rapid-cycle testing to assess specific design choices. Larger RCTs should come later, once the intervention is stable, implementation is consistent, and there is initial evidence that target teacher and student outcomes are improving.
2. Ask how it works before asking whether it works
Before asking whether a tool improves student achievement, we should ask whether it changes the behaviors it is designed to change. For example, does an AI coaching tool actually shift how teachers plan, instruct, or respond to student thinking? If not, we should focus on addressing those issues first, before moving to assess student impact.
3. Let the questions drive the method
AI tools make it possible to run frequent, low-cost experiments and collect detailed usage data. The field should leverage these capabilities through A/B testing and continuous evaluation, rather than relying solely on methods designed for static programs. The design and methodology should be made to match the research question, not the other way around.
Why this matters for decisionmakers
The shift that we are proposing has implications for system leaders, practitioners, researchers, and funders. A state investing in AI today cannot wait years for evidence from a study on a version of an intervention that no longer exists. Evidence that arrives too late cannot guide decisions about adoption or scale—yet many organizations reach significant scale without even basic outcome data on teaching and learning.
Implementation R&D offers a more practical path to ensuring that tools are effective before they are taken to scale. It enables organizations to generate credible, early evidence, refine their approaches, and build toward more rigorous evaluation.
The Research Partnership for Professional Learning’s (RPPL’s) Shared Measures Toolkit is one example of implementation R&D in practice. Rather than beginning with a fully developed measurement product, RPPL worked with professional learning organizations and researchers to identify the aspects of high-quality instructional materials and curriculum-based professional learning implementation that practitioners most needed to understand and improve. The resulting measures are now being tested, refined, and validated across multiple contexts to determine whether they are psychometrically sound and whether they generate useful information for continuous improvement and decisionmaking. This work is building a measurement infrastructure that can help districts and professional-learning organizations better understand implementation quality, compare results across contexts, and identify high-leverage practices associated with stronger teacher and student outcomes.
This approach strengthens the conditions for eventual RCTs by sequencing implementation R&D and causal testing mechanisms without lowering standards. If we want to ensure AI tools contribute to the improvement of teaching and learning, we need to study how those AI interventions are deployed in near-to-real time, using iterative experimentation to move from early design to long-term impact.
AI tools should be held to high standards. But to make those standards achievable, the education sector must move toward a new paradigm for research focused on producing iterative evidence and proof of impact on the outcomes that matter most.
-
Acknowledgements and disclosures
Thank you to Dr. Sarah Johnson and the Teaching Lab team for being in conversation with us and sparking the idea for this piece.
The Brookings Institution is committed to quality, independence, and impact.
We are supported by a diverse array of funders. In line with our values and policies, each Brookings publication represents the sole views of its author(s).
Commentary
AI is rapidly changing education and research needs to keep up
July 21, 2026