Can an AI Tutor Help Middle Schoolers Learn Math? What a Study of Goblins Among 6,000+ Students Found

At a time when the debate over AI and screen time is dominating education headlines, there is remarkably little independent evidence about whether these tools help students learn. The first report from Accelerate’s Call for Effective Technology (CET) grant program offers early and encouraging findings for Goblins, a math platform providing interactive AI support to students. 

In Arizona’s Deer Valley Unified School District, middle school students who solved at least two problems a week on Goblins scored about 0.11 standard deviations (SDs) higher on MAP math assessments than similar students who had little or no usage, equivalent to nearly three months of additional learning. It’s a promising, statistically significant finding, but not yet definitive proof that Goblins alone caused the difference. To help answer that question, Accelerate is working to launch a randomized controlled trial with Goblins during the 2026–27 school year.

What was the study?

Researchers from the Center for Educational Data Science and Innovation at the University of Maryland conducted a quasi-experimental study (QED) among students in grades 6–8 in Deer Valley during the 2025–26 school year. The analysis included 6,065 students, a large majority of the suburban district’s middle schoolers. 

Students were not randomly assigned to use Goblins. The study compared students who solved an average of at least two problems a week with those who solved fewer or none at all. Researchers statistically adjusted the groups to account for prior achievement and other observable student characteristics such as race, disability, and English-learner status, along with their teachers’ past performance in raising student achievement. But this approach cannot account for differences that the researchers could not observe.

What did the study find?

Let’s start with the headline. Students who solved at least two problems a week on Goblins achieved nearly three months of additional math learning (0.11 SDs higher on MAP assessments) compared with similar district peers with little or no usage, after adjusting for achievement levels at the start of the school year.

This is a notable, statistically significant finding. Still, students and teachers chose whether and how much to use Goblins, so other differences, such as student motivation and teacher decisions, may have contributed.

The researchers also looked at outcomes by usage level and found that higher usage beyond the two-problems-per-week threshold was associated with larger gains. Students who solved at least 20 problems per week had gains equivalent to roughly one additional year of learning. However, only 164 students (2.7%) reached that level of usage. We don’t know whether that intensity is reachable at scale or whether the small group of students who achieved that usage level differs from peers in ways beyond how much they used the platform. 

How confident should we be?
  • Heavier users may have differed in ways the study couldn’t measure. Because students were not randomly assigned, the results are open to selection bias: factors such as home support could explain part of the gains. Researchers statistically adjusted the groups to look similar based on available characteristics, but they couldn’t account for what wasn’t measured.
  • The groups weren’t fully comparable at the start. Even after weighting, they differed in starting achievement levels and disability status. Researchers adjusted for these gaps, and the study met What Works Clearinghouse baseline-equivalence standards. Still, these differences make it harder to know how much of the result reflects Goblins usage rather than baseline differences between the students.
  • The study answers a narrower question than whether offering Goblins improves learning. Rather than comparing students offered Goblins with those who weren’t (an intent-to-treat analysis), it defined groups by usage levels. As a result, the study does not answer the question: “What is the impact of providing Goblins access to students?” Instead, it answers: “How did outcomes differ between students who used Goblins more and those who used it less?”
Who participated in the study?

Of the district’s 7,083 sixth- through eighth-grade students, 6,065 were included in the analytic sample and given access to Goblins. The sample was large and aligned with district demographics — predominantly white, with a significant Hispanic student population and relatively few Black or Asian students.

One gap: researchers could not analyze results for economically disadvantaged students because data on free and reduced-price meal eligibility was missing for about one in five students.

The researchers also looked at how results varied across student groups and saw possible larger benefits for students with IEPs, Hispanic students, and English learners, but these subgroup results are preliminary.

Future studies should test these patterns further, examine results for economically disadvantaged students, and include districts with different demographics and larger shares of Black and Asian students to better understand how Goblins works across diverse student populations.

What did implementation look like?

On average, students used Goblins for 5.3 hours over the school year (about 27 weeks), including roughly 3.5 hours of teacher-assigned work and 1.8 hours of independent practice. But usage varied widely. Just over half of students (3,088) met the research team’s threshold of two problems solved per week. About two-thirds (67%) of students averaged fewer than 10 minutes a week on the platform, while 15% averaged at least 20 minutes.

Teachers chose whether and how to use Goblins in their classrooms and were not given a specific usage target. The district recommended that Goblins be used for supplemental rather than core instruction, but implementation varied. Usage records can’t distinguish whole-class from targeted use, but students who activated a Goblins account started the year about 15 percentile points behind their classmates in math, suggesting that teachers may have used Goblins primarily for Tier 2 and Tier 3 instruction. 

These patterns offer lessons for how Goblins might be implemented and further studied in the future. Most usage came through teacher-assigned work, reinforcing the importance of clear expectations and dedicated instructional time. Five problems per week may be a reasonable minimum target to test in future implementations, with enough time scheduled during Tier 2 or Tier 3 instruction to reach it.

What comes next?

A well-designed randomized trial is the natural next step. The good news is that we are already working toward one, with Goblins continuing as a CET grantee alongside our second cohort.

This study has given us a promising finding. We’re grateful to Goblins and Deer Valley for partnering with us on this research and committing to share the findings publicly. That reflects a shared north star of improving learning and an understanding that honest evidence is the surest path to better outcomes for students.

This is the beginning of an evidence-building process, not the end. Stronger, more rigorous evidence will provide a clearer understanding of how to achieve consistent results.


About the Call for Effective Technology (CET)

Through our Call for Effective Technology (CET) grantmaking program, Accelerate commissioned independent evaluations of AI-powered educational tools during the 2025–26 school year. The studies in our first CET cohort collectively involved roughly 21,000 students across 14 states, tested tools under real classroom conditions, and used quasi-experimental designs. Our second CET cohort includes 11 grantees, with several running randomized controlled trials.

These tools are evolving faster than typical research timelines, so we are sharing results as quickly as possible. Timely evidence helps providers identify what to improve and gives education leaders better information for their purchasing decisions. We’re grateful to the providers and districts that took part and committed to sharing the findings publicly, whether they are encouraging, disappointing, or inconclusive. These studies will not definitively settle the debate about AI in classrooms, but they represent an important step toward understanding which tools help students learn and under what conditions.

Stay Connected!

Learn What’s Working for Students

Get the latest from Accelerate, including news, insights, and updates on our work.

Subscribe
Captcha