Back to Queen Creek

Is AI Helping Our Children? Arizona Is Running the Experiment Live

An Arizona charter school teaches core academics in two hours a day through AI software, with guides instead of teachers. A randomized trial suggests everything depends on whether the tool hints or hands over the answer.

Pierce Keller

August 20, 20265 min read

Bar chart of results from a randomized trial of GPT-4 in high school mathematics. During practice, students using a standard chatbot scored 48 percent higher than classmates with no AI access and students using a hints-only tutor scored 127 percent higher. On an exam with the tool removed, the standard-chatbot group scored 17 percent lower, while the hints-only group showed no significant change. Source: Bastani et al., PNAS, 2025.
Bar chart of results from a randomized trial of GPT-4 in high school mathematics. During practice, students using a standard chatbot scored 48 percent higher than classmates with no AI access and students using a hints-only tutor scored 127 percent higher. On an exam with the tool removed, the standard-chatbot group scored 17 percent lower, while the hints-only group showed no significant change. Source: Bastani et al., PNAS, 2025.

Most states are still writing guidance on artificial intelligence in schools. Arizona approved a school built on it.

Unbound Academy, an online only charter serving grades four through eight, was cleared by the Arizona State Board for Charter Schools on a 4 to 3 vote. Its model compresses core academics into roughly two hours a day, from 8:30 to 11:00 in a school day that runs 8 to 2, delivered through AI powered educational software.

The rest of the day goes to workshops in things like financial literacy, public speaking and entrepreneurship.

It does not employ classroom teachers in the usual sense. It employs guides, who coach students through the software, step in when someone is stuck, and run the afternoon sessions. The approach comes from Alpha School, a private school in Austin, Texas, that has used it since 2016. Unbound is open and enrolling now.

Whatever one makes of it, it is the most direct test in the country of a question the research has just started to answer properly. And the research says the details decide everything.

The trial that split the difference

Researchers gave roughly a thousand high school students access to GPT-4 during mathematics practice. Practice scores rose 48 percent against students working without it.

Then the tool was withdrawn for an exam. Those students scored 17 percent worse than classmates who had never used it.

The study, published in the Proceedings of the National Academy of Sciences, found the students had used the model as a crutch. What makes the result genuinely difficult is that nothing looked wrong while it was going on. Practice scores, homework quality and engagement all improved. The deficit appeared only once the tool was absent.

The configuration that changed the outcome

A third group used the identical GPT-4 model, instructed to supply incremental hints and never the full answer.

They improved the most during practice, by 127 percent, and showed no significant harm on the exam afterward. Same model, same students, same material. The only difference was whether the software would hand over an answer.

That is the finding to hold against any AI driven school model, including one operating in the East Valley. The question is not whether software can teach. It is whether the software is built to make the student do the thinking, and whether an adult is checking that it does.

What the supporting evidence suggests about design

The strongest positive result available happens to point the same direction.

A World Bank randomized trial in Benin City, Nigeria, put first year senior secondary students through six weeks of after school English sessions using Microsoft Copilot, supervised by teachers. The gain was 0.31 standard deviations across the full assessment, which covered English, AI knowledge and digital skills, and 0.23 on English alone, the study's main outcome.

The researchers report the program outperformed 80 percent of the interventions in a comparison database of randomized trials in developing countries.

The two features that stand out are that the sessions were structured and that a teacher supervised them. That is much closer to the hint giving arm of the Turkish trial than to a child working alone.

The effects were also uneven, with the largest gains among female students and among those who arrived with higher initial academic performance, which is a caution for any model that assumes independent work.

The adults are still catching up

Gallup and the Walton Family Foundation found roughly six in ten teachers used AI in the past school year, with about a third using it weekly. Weekly users report saving an average of 5.9 hours a week, close to six weeks over a school year.

Only 18 percent said they had received formal guidance from administrators on how the tools should be used. About a third said they had received none at all.

Catching misuse is not a workable plan

Schools hoping to manage this through detection have a problem. Stanford researchers tested widely used AI detectors on essays written by human non-native English speakers and found they consistently misclassified that writing as AI generated, flagging more than half of the non-native essays while performing nearly perfectly on writing by native speaking US eighth graders.

The tools were reading plainer sentence construction as machine output. Several universities have since disabled them: Vanderbilt switched Turnitin's AI detector off in August 2023, citing reliability, false positives and the effect on international students.

What has not been shown

This is one subject in one country across four sessions, covering about 15 percent of the semester's mathematics curriculum. Math is the easiest subject to build a hint giving tutor for, because problems have fixed answers and the errors students make are already mapped. No comparable trial exists for writing or history.

Every result here is measured in weeks, not years, so the durability of the effect is unknown in both directions.

None of this explains falling national test scores either. Those declines set in after 2012 among the lowest performing students, years before ChatGPT, and the timeline does not support blaming AI for them.

The question worth asking any school

The findings do not make a case for banning these tools, and they do not make a case for turning a child loose with one. They make a case for a single question, which works whether a school is fully AI driven or barely touching it:

When a student is stuck, does the software give the answer, or the next hint?

In the only controlled comparison available, that was the whole difference between learning and not learning. It is a fair thing to ask any principal, and a fairer thing to ask one whose model depends on the answer.

Sources

  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. "Generative AI without guardrails can harm learning: Evidence from high school mathematics." Proceedings of the National Academy of Sciences, 2025.
  • Gallup and the Walton Family Foundation, teacher AI surveys, 2025.
  • AcadeResearch, "Practice Up 48%, Exams Down 17%: Does AI Stop Children From Learning?", August 2026.
  • Arizona State Board for Charter Schools, approval of Unbound Academy; Unbound Academy school site.
  • Liang, W. et al. "GPT detectors are biased against non-native English writers." Patterns, 2023.
  • World Bank Education Global Department, "From Chalkboards to Chatbots," Nigeria.
Share

Pierce Keller

Pierce Keller writes about community life, schools, public safety, and local events in Queen Creek.

Related Stories

More in Arizona