Loading summary
A
If this episode makes you think, please let us know in the comments and support us by subscribing and leaving a review. Thank you. Today we are exploring the critical challenge of evaluating AI based tools in education, drawing insights from an illuminating piece titled AI is Rapidly Changing Education and Research Needs to Keep Up. This article, published on July 21, 2026 by Stacy Alicea and Megan McCormick, dives into why our traditional methods for gauge and the effectiveness of new educational technology just aren't cutting it for the fast paced world of artificial intelligence. And what they reveal is quite striking. Nearly 2/3 of teachers now report that they use AI in their work, yet only about 1 in 5 edtech products and classrooms today actually has any evidence that they can improve teaching and learning outcomes. That's a huge gap, isn't it? Now, first, let's talk about the problem itself. It's a fantastic starting point. The conversations about AI in education often begin with the right question Is there evidence of the effects of this technology? Of course we want to know what impact an intervention will have before it lands in a classroom, impacting students and teachers. But as Alicea and McCormick point out, the next question often becomes too narrow. We tend to jump straight to what have we learned about this technology from randomized controlled trials? And the issue is randomized controlled trials, or RCTs, are notoriously difficult to conduct in education. They take a lot of time, a lot of resources, and they typically require a stable, well defined intervention that doesn't change much. That's where the pursuit of evidence often stops, because RCTs are just not readily available for the vast majority of new tools. The authors make a really compelling case that AI is a fundamentally different type of intervention from what RCTs were designed to evaluate. Think about it. An AI tool today, let's say one that helps teachers with instructional feedback or analyzes student work, might look completely different next month. Developers are constantly updating models, refining prompts, and the ways teachers and students interact with these tools are evolving in real time. It's an evolution, not a revolution, yes, but a very fast evolution. An RCT assumes that the thing you're testing stays pretty much the same throughout the study, but with AI that assumption just doesn't hold. So while R ETS have been essential for building evidence in areas like class size reductions or the science of reading, they simply can't keep pace with the rapid innovation in AI. We need strong evidence now, not perfect evidence years later, especially when school leaders are making decisions about adopting AI tools in their systems. This is why we need to rethink our approach to AI education research. The second big idea here is that instead of the traditional RCT, Aliceer and McCormick advocate for something called Implementation Research and Development Tool or Implementation R and D. This approach is much better suited to the evidence needs of AI technology because it's traditionally used for studying early stage products. It's all about rapid test and feedback and refinement. So it helps us understand how a tool is designed, delivered and used in real world settings before we even think about a massive large scale evaluation. It gives us that initial evidence on a product's effects on teaching and learning. And crucially, if a product isn't working, it helps us figure out why. Now the authors are Implementation R and D isn't meant to replace RCTs entirely. Rather, it's the foundation upon which to build research that eventually proves cause and effect. It focuses on learning what works before testing whether it works at scale. I think that's a really important distinction and it connects directly to the purpose over technology pillar of my core philosophy. We're not just asking if the tech works, we're asking if it serves its educational purpose and how it's doing that. What's fascinating is that some national efforts are already embracing this. The Institute of Education Sciences, or ies, for example, has funded generative AI R and D centers that prioritize iterative development and pilot testing before any formal trials. And then you have partnerships like Lean Lab Education and Boston University's EVAL Initiative, both supporting technology through these cycles of real world testing and refinement. This iterative feedback loop is really powerful. This brings us to a really practical framework of three guiding principles that Aliceer and McCormick offer for building evidence for AI tools. The first principle is to set build evidence in stages for early stage AI tools. The focus should be on feasibility and rapid cycle testing to assess specific design choices. Think of it like a sandbox, right? You're experimenting, trying things out, seeing what sticks. You're trying to prove the concept works. The larger, more resource intensive RCTs should come much later, only once the intervention is stable, the implementation is consistent, and there's already some initial evidence that target teacher and student outcomes are improving. This aligns beautifully with my act now principle within the seven Lessons for AI Adoption, which encourages educators not to wait for perfect conditions, but to get started with hands on pilots and testing. So you can't expect a polished final product if you're not allowing for early messy iteration. The second principle really resonates with me. Ask how it works before asking whether it works. This is crucial for any educational intervention, not just AI tools. Before we even consider if an AI tool improves student achievement, we should be asking if it's actually changing the behaviors it was designed to change. For example, if you're using an AI coaching tool, is it genuinely shifting how teachers plan lessons, how they instruct, or how they respond to student thinking? If those foundational changes aren't happening, then we need to address those issues first. There's no point in measuring student impact if the tool isn't even moving the needle on teacher practice. It's about understanding the process and the productive struggle, not not just the eventual outcome. The real value is not in what the machine produces, but in how the student, or in this case, the teacher, responds and adapts. And the third principle is to let the research questions drive the method. AI tools give us incredible capabilities to run frequent low cost experiments and collect really detailed usage data in real time. The field should be leveraging these strengths through things like AB testing and continuous evaluation, instead of clinging solely to methods that were designed for static, unchanging programs. An ongoing study of an AI coaching tool from Teaching Lab, for instance, compares a standard version of the tool to one that emphasizes student centered instruction. This allows them to see if subtle design differences actually shift coaching practices. Because the tool captures interactions in real time, researchers can analyze responses to feedback as they occur, enabling faster, more responsive research. This allows researchers to ask much more nuanced questions rather than just a blunt does it work or not? It's about precision in evaluating AI tool effectiveness. So what does this mean for school leaders? And really for every educator listening who is thinking about implementing AI in classrooms? The shift that Alassia and McCormick are proposing has massive implications for decision makers. A state or a district investing in AI today cannot afford to wait years for evidence from a study on a version of an intervention that will likely no longer exist by the time the data comes in. Evidence that arrives too late simply cannot guide decisions about adoption or scaling up. Sadly, many organizations reach significant scale without even basic outcome data on teaching and learning. That's a real risk. Implementation R and D offers a much more practical path to making sure that the tools we're using are actually effective before we take them to scale. It empowers organizations to generate credible early evidence, allowing them to refine their approaches and and build towards more rigorous evaluation down the line. It's about starting with why, not how? And continuously checking that your how is serving your why. An excellent example of this Implementation R and D in practice is the Research Partnership for Professional Learning's Shared Measures Toolkit. Instead of starting with a fully developed measurement product, the RPPL worked collaboratively with professional learning organizations and researchers. They identified what aspects of high quality instructional materials and curriculum based professional learning practitioners really needed to understand and improve. The resulting measures are now being tested, refined and validated across multiple contexts, ensuring they're psychometrically sound and that they generate useful information for continuous improvement and decision making. This kind of work is building a crucial measurement infrastructure that can help districts and professional learning organizations better understand implementation quality, compare results, and pinpoint high leverage practices that lead to stronger teacher and student outcomes. This type of iterative evaluation is critical for responsible AI edtech evaluation. I think this approach strengthens the conditions for eventual rct. It doesn't lower the bar, it just sequences implementation R and D before the causal testing mechanisms. It's an evolution, not a revolution, in how we do education research. If we truly want AI tools to genuinely improve teaching and learning, we absolutely need to study how those AI interventions are being deployed in near to real time, using iterative experimentation to move from early design to long term impact. This way, we're teaching students not to outsmart machines, but to outthink them and ensuring the tools we use truly serve their learning. The core message is clear, if we want better tools, we need better, faster research. That's all for today. Thanks for listening.
Episode: Rethinking Edtech Evaluation
Date: July 28, 2026
Host: Dan Fitzpatrick, The AI Educator
In this episode, Dan Fitzpatrick delves into the urgent challenge of evaluating AI-based educational tools, drawing insights from the recent article, “AI is Rapidly Changing Education and Research Needs to Keep Up” by Stacy Alicea and Megan McCormick (published July 21, 2026). The episode explores the limitations of traditional evaluation methods like randomized controlled trials (RCTs) in the era of ever-evolving AI and advocates for a more adaptive approach, such as Implementation Research and Development (Implementation R&D). Listeners are guided through the practical implications of this shift for educators, school leaders, and edtech decision-makers.
“RCTs are notoriously difficult to conduct in education… an AI tool today… might look completely different next month. Developers are constantly updating models.” — Dan Fitzpatrick (02:40)
“It’s an evolution, not a revolution, yes, but a very fast evolution.” — Dan Fitzpatrick (03:30)
“We're not just asking if the tech works, we're asking if it serves its educational purpose and how it's doing that.” — Dan Fitzpatrick (06:40)
“You can't expect a polished final product if you're not allowing for early messy iteration.” — Dan Fitzpatrick (11:20)
“There’s no point in measuring student impact if the tool isn’t even moving the needle on teacher practice.” — Dan Fitzpatrick (13:00)
“It’s about precision in evaluating AI tool effectiveness.” — Dan Fitzpatrick (14:45)
“Evidence that arrives too late simply cannot guide decisions about adoption or scaling up.” — Dan Fitzpatrick (15:45)
“This type of iterative evaluation is critical for responsible AI edtech evaluation.” — Dan Fitzpatrick (18:10)
“If we want better tools, we need better, faster research.” — Dan Fitzpatrick (19:30)
By reframing edtech evaluation for the AI era, this episode empowers educators and decision-makers to pursue more agile, relevant, and actionable research—ensuring technology always serves learning, not the other way around.