In 2020 - years before large language models became publicly available - I argued that schools may want to consider radically reorganizing the academic day. Specifically, I proposed short, intensive academic sessions in the morning followed by extended afternoon blocks of focused work.
My reasoning had nothing to do with artificial intelligence or digital tutors – rather, it stemmed entirely from what we know about the biology of learning: the brain's changing capacity for sustained cognitive work across the day and the conditions under which newly acquired knowledge is most effectively consolidated and adapted.
At first glance, this schedule looks remarkably similar to what several AI-powered schools now advertise. But the resemblance is almost entirely superficial.
My proposal retained a broad academic curriculum (including languages, the arts, civics, and history) and used the extended afternoon blocks to deepen the morning’s learning through laboratories, revision, and sustained projects. The goal was not to replace explicit instruction with entrepreneurship or life-skills training, but to strengthen transfer by guiding students through the application of newly acquired knowledge in increasingly complex contexts.
Whether my proposal is ultimately optimal is almost beside the point. The important thing is that there are sound theoretical reasons to expect that reorganizing the school day, by itself, could improve learning outcomes. And that creates a non-trivial scientific problem for AI schools:
If reorganizing the school day can improve learning...
…and if AI schools simultaneously reorganize the school day while introducing AI tutoring…
…how can we confidently attribute any reported improvements to the AI?
Welcome to the Confound Dilemma.
CONFOUNDS ABOUND
In a well-designed experiment, researchers attempt to isolate a single variable. Everything else in the environment is held as constant as possible so that any observed change can reasonably be attributed to the variable being tested.
Once research enters the real world, however, it becomes increasingly difficult to control every meaningful variable. Schools, hospitals, businesses, and communities are constantly changing in countless ways, making it difficult (if not impossible) to determine which changes actually produce any observed outcome.
These additional variables, which change alongside the intervention of interest, are known as confounds – and they provide alternative explanations for any observed outcome which make causal conclusions far more uncertain.
Imagine a city installs speed cameras, lowers speed limits, repaves dangerous intersections, launches a road-safety campaign, and hires more police officers - all at the same time. Now imagine one year later, traffic fatalities decline.
Which intervention caused the improvement?
The honest answer is: we don’t know. Any one of those changes could have helped; several may have worked together; one may have offset the effects of another. Without isolating each intervention, attributing the entire improvement to, say, speed cameras alone would be scientifically unjustified.
Now replace ‘speed cameras with ‘AI Tutoring’, and you have the central problem facing today’s digital-first schools.
THE ATTRIBUTION PROBLEM
When several emerging school systems introduced AI-powered instruction, they didn’t simply add a new tool to an otherwise unchanged educational model. They simultaneously transformed nearly every aspect of schooling: the schedule, the role of teachers, the curriculum, assessment practices, parental expectations, mentoring practices, and much more.
Recently, these schools have reported remarkable improvements on measures of student learning. Importantly, these schools largely attribute these gains almost exclusively to the digital programs: their central claim is that replacing traditional teachers with AI tutors improves learning.
In previous posts, I’ve argued that many of these reported educational gains are likely overstated and may not reflect genuinely deeper, transferable learning.
That debate, however, is not the focus here. So, for the remainder of this article, let’s make the generous assumption that every reported improvement in these schools is genuine. Even under that assumption, a completely separate scientific concern remains: attribution.
How do we know the AI deserves the credit? When an entire system changes at once, every change becomes a plausible explanation for any subsequent improvement.
Again, I’ll make what I consider to be a very generous assumption: that students in these schools are genuinely learning more. Even so, there remain dozens - if not hundreds - of potential confounds. Luckily, we needn’t identify every single one: if just a handful provide credible alternative explanations for reported gains, then attributing those gains to AI becomes scientifically unjustified.
Let’s examine just three.
#1 ORGANIZATION
As noted above, there’s a strong argument to be made that reorganizing the school day could improve learning. This argument has nothing to do with artificial intelligence – rather, it stems from two basic principles of neuroscience.
The first is the brain’s limited capacity for sustained, effortful learning.
Acquiring new knowledge is metabolically demanding: as students engage in prolonged periods of focused cognitive work, the brain’s capacity for efficient learning diminishes . Although the precise biological mechanisms remain under investigation, one leading hypothesis implicates the depletion of astrocytic glycogen (the brain’s readily available reserve of metabolic energy).
Together with broader research on attention and distributed practice, this provides a strong theoretical rationale for structuring periods of intensive learning around strategically timed breaks rather than simply extending lessons indefinitely.
The second principle concerns flow-state.
Once new knowledge has been consolidated, the goal shifts from acquisition to application. Here, the opposite pattern emerges: rather than short, intensive periods of work, learners benefit from long, uninterrupted periods in which they can revise, experiment, and apply newly learned ideas to increasingly challenging contexts. These are the conditions under which flow-state is most likely to emerge - a neurological state associated with heightened focus, sustained engagement, and skill refinement.
Whether this specific organizational model is optimal remains an open question – but the point I’m trying to make is that there are sound scientific reasons to expect that reorganizing the school day, by itself, would improve learning.
Now consider what has happened in many AI schools: they have both introduced digital tutoring while simultaneously reorganizing the school day. If learning subsequently improves, how would we know which intervention deserves the credit?
Now consider the counterfactual: what if we implemented this same organizational change while retaining high-quality human teachers? If comparable gains followed, it would suggest that reorganizing the day – not the AI - was the primary driver of improvement.
#2 HIDDEN PREPARATION
One of the most impressive claims made by AI-powered schools concerns their Advanced Placement (AP) results.
Several schools have reported that their students consistently earn top marks in these exams – results which, if genuine, would appear to provide compelling evidence that AI tutoring can replace traditional instruction.
Even setting aside the obvious and well-documented issue of AP grade inflation, a far more fundamental problem remains for AI-schools.
Anyone who has taught an AP course knows that students cannot master an entire AP curriculum through thirty minutes a day of generic tutoring focused primarily on mathematics, reading, and basic science. The sheer breadth and specificity of AP content makes that impossible.
So, assuming these reported AP results are accurate, how are they happening?
The explanation becomes clearer when we look more closely at how several prominent AI schools prepare students for these exams. One publicly reports that students complete approximately 75 hours of dedicated AP preparation before sitting each exam. Another schedules a two-week intensive AP preparation bootcamp immediately prior to the examination period.
Those details change the interpretation entirely.
The 75 hours of explicit practice and dedicated two-week prep program constitute additional educational instruction. They are targeted, assessment-specific intervention. Once we recognize this, exceptional AP performance can no longer be attributed solely (if at all) to the school’s everyday AI tutoring model.
If those same 75 hours (or the same two-week bootcamp) were delivered within a conventional school by experienced teachers, we simply don’t know whether comparable – or even better - results would follow.
#3 PARENTAL INVOLVEMENT
Consider the type of person who would enroll their child into an AI-powered school.
First, these schools are prohibitively expensive, meaning they’re financially accessible only to a relatively small segment of the population. Second, parents must already be sufficiently dissatisfied with their child’s current schooling to actively seek an alternative. Third, they must invest time researching an entirely new educational model. Finally, many families enrolled in AI-schools undertake lengthy daily commutes (or even relocate) to make attendance possible.
Do these sound like the actions of indifferent parents – or of parents deeply committed to their children’s education?
It should come as little surprise that decades of educational research consistently associate this kind of positive parental involvement with stronger academic outcomes. Although there are important caveats (such as parents becoming overly controlling, creating unhealthy pressure, or undermining teachers), children generally benefit when parents are actively engaged in their education. Such involvement helps cultivate stronger academic self-concepts, higher expectations for success, greater engagement with school, and more productive partnerships between students and teachers.
One might argue that these parents were probably just as involved before enrolling their children in an AI school. But that’s beside the point. Families who choose these schools are systematically different from the average family attending a conventional school.
They have actively sought out an alternative educational philosophy, invested significant time and money to access it, and aligned themselves with a model they believe will benefit their child. Moreover, the very act of choosing such a school is itself likely to increase parental engagement, as families become more invested in reinforcing a model they have deliberately selected.
And once again, we’ve encountered a confound.
If students in these schools perform well, how much of that success reflects the AI tutor - and how much reflects an unusually committed family?
Now imagine if we fostered this same level of parental commitment within a traditional school with experienced teachers? Again, if comparable gains followed, parental engagement (not AI) would emerge as the more plausible explanation. And if even greater gains followed, we’d have genuine reason to question whether the AI was helping at all.
SO NOW THEN…
Over the past few years, an increasing number of researchers and educators have called for rigorous, independent evaluations of AI-powered schools.
The reason isn’t that they oppose innovation – in fact, quite the opposite. If AI schools have genuinely discovered a better way to educate children, we should understand exactly what is happening so those benefits can be extended to every school.
But that requires more than carefully crafted press releases and vague explanations of how the model functions.
For the sake of argument, I’ve assumed here that every reported gain is genuine. Even under that generous assumption, we’ve identified three plausible confounds (and there are doubtless many more) that could provide alternative explanations for student improvements. Until these factors are isolated and evaluated independently, we simply cannot know how much gain should be attributed to AI, how much belongs to the broader school model, or which elements are actually worth replicating elsewhere.
That is why rigorous comparative research matters.
If AI tutors truly improve learning, well-designed studies will demonstrate it (and, as discussed elsewhere, there are specific contexts in which they appear to). But if other elements of these schools prove more important than the AI itself, we need to know that as well. Either outcome will represent meaningful progress.
My goal here (as elsewhere) has never been to attack technology – I aim only to better understand what actually helps students develop. I am not anti-tech; I am pro-learning.
Luckily, science doesn’t require we trust claims – it demands we test them.




An article where I agree with everything!
I'm also a big fan of reorganizing the school day similar to how you've described.
I'll add one more big confound to your list that often appears in these "AI" schools.
Assuming a lot of this article was written with Alpha School in mind. Alpha also runs its computers closer to a "kiosk". No YouTube, no multitasking, no LLM chatbots (by default). What's actually more "revolutionary" than the AI itself is just be effective control of the operating system. (Guided Access)
The way we use tech in schools is like covering the students desks with marshmallows and telling them not to eat them, the students willpower rapidly depletes knowing Youtube is 2 second click away. Every lesson with an open operating system PC becoming a self-control contest.
OS-level restriction (Guided Access) is itself an unmeasured confound sitting inside many AI school's bundle of simultaneous changes, and it's probably doing more work than the AI.
I'm sure you're familiar with this in in motivation/choice psychology, like how "3" choices is a great number. How many colour options does Apple give for their flagship iPhones? Usually 3. Or like the marketing research selling Jam at a stall, where having too many options resulted in less engagement/sales. Fewer options (but not zero) giving the most engagement.
Imagine instead that when you get on the computer, the interaction space is so confined that it has a similar number of choices to paper.
• make origami/planes
• draw sketches
• do the lesson
Comapred to computers where you get near infinite choice. Imagine instead you get on the computer, log in, then you're presented with 3 options of what to work on for the next 25 minutes in blocks of explicit pomodoro computer instruction:
• Geometry
• History
• Chemistry
The only "AI" going on is the knowledge graph and spaced repetition math going on behind the scenes to dictate which of those three options is best to present to the student (Which is barely AI, as it was invented so long ago and is often just a simple algorithm of a few lines).
The student isn't able to screenshot a question and ask AI to answer it, because screenshots are disabled, browsers are disabled, everything is disabled except the lesson. Pressing the home button only leads back to the 3 options. Ctrl+Alt+Delete and other system commands are disabled. If they need a browser, the teacher can enable it for short time blocks (10 minutes of Google, then it disables itself again).
This is how some of these "AI" schools are approaching tech. You can remove the AI and it'll still be far better than regular tech use in school.
Sadly I've literally not been able to find a study that has DIRECTLY compared computers with guided access vs without guided access with every other variable held constant. That would probably be one of the most important studies for this topic. I suspect a study like that would help explain a good chunk of the credit towards the "AI" schools.
Anyway, great article Jared :)
The marshmallow desk experiment is a perfect frame — we keep designing environments that demand willpower from people we've already depleted with context switches and notification pings. Adding AI on top of that without redesigning for recovery and focus time isn't productivity. It's accelerated burnout with better charts.