AI Can Improve Student Performance. But Are Students Learning?


Generative AI can help a student write a better essay, solve a difficult problem, improve an explanation, or complete an assignment faster.

But there is a question educators need to ask more often:

Did the student actually learn anything?

The OECD Digital Education Outlook 2026 raises an important distinction between performance and learning.

Emerging evidence reviewed by the OECD suggests that access to general-purpose generative AI can improve the quality of students’ work while they are using it. But better AI-assisted performance does not automatically translate into better independent performance later.

That distinction matters enormously.

A student can produce an excellent answer with AI while developing surprisingly little ability to produce, explain, or evaluate that answer independently.


The illusion of understanding

Imagine two students trying to solve the same physics problem.

The first student struggles.

They draw a diagram. They choose the wrong equation. They notice something does not make sense. They go back. They remember something their teacher explained last week. They try again.

Eventually, they solve it.

The second student gives the problem to an AI assistant.

Within seconds, the AI identifies the relevant equation, substitutes the values and produces a beautifully structured explanation.

Who performed better?

Probably the second student.

But who learned more?

That is a completely different question.

The uncomfortable possibility is that educational technology can optimise the thing we can easily see — the finished answer — while quietly removing some of the cognitive activity that produced learning in the first place.


Struggle is not always a problem to eliminate

As teachers, we naturally want to help students when they struggle.

But not every struggle is harmful.

Sometimes the moment when a learner thinks:

“Wait… why doesn’t this work?”

is precisely the moment when learning begins.

Trying to remember something strengthens retrieval.

Recognising an incorrect assumption develops metacognition.

Comparing two possible solutions develops reasoning.

Explaining why an answer is correct forces the learner to organise their understanding.

Getting something wrong and correcting it can expose a misconception that would otherwise remain hidden.

If an AI system immediately removes all of these moments, it may make learning feel easier while also removing some of the processes through which understanding develops.

Recent research has increasingly described this challenge in terms of preserving productive struggle: difficulty that is challenging enough to require thought, but supported enough that the learner can continue progressing.

This changes how I think educational AI should be designed.


Perhaps the best AI tutor should sometimes refuse to answer

Most general-purpose AI assistants are optimised to be helpful.

Ask a question and they answer it.

Ask for an explanation and they explain it.

Ask them to solve something and they solve it.

That behaviour makes perfect sense for a general assistant.

But a learning system has a different objective.

Its job is not necessarily to help the learner finish the current problem as quickly as possible.

Its job is to help the learner become capable of solving future problems without needing the system.

Those objectives can conflict.

Imagine a student enters:

“Calculate the acceleration. Just give me the answer.”

A conventional assistant might calculate it immediately.

A learning-oriented AI might instead respond:

“Before we calculate it, which two quantities do you think we need?”

If the learner cannot answer, the system could provide a small hint.

If they are still stuck, it could reveal another step.

Only when necessary would it provide the complete explanation.

The AI is still helping.

But it is deliberately preserving some of the thinking for the learner.


This is where active recall becomes especially important

There is another reason I find this question fascinating.

One of the most powerful principles in learning science is retrieval practice.

Reading an explanation can create a feeling of familiarity.

But being asked to retrieve an idea forces us to discover whether that knowledge is actually available to us.

AI gives us an extraordinary opportunity to make this process adaptive.

Instead of continuously generating explanations, an educational AI could ask:

“Explain that back to me in your own words.”

“Why did you choose that equation?”

“What would happen if this variable doubled?”

“You made this mistake earlier. Can you spot what changed this time?”

“Try the same concept again without my help.”

Now AI is not simply delivering knowledge.

It is creating opportunities for the learner to retrieve, explain, apply and test that knowledge.


We may need a different measure of AI success

This also suggests that EdTech companies may be measuring the wrong things.

Imagine an AI tutoring system reports:

  • 94% of questions successfully answered
  • homework completed 35% faster
  • students required fewer attempts
  • satisfaction increased significantly

Those numbers sound fantastic.

But I would want to know something else.

What happens when the AI disappears?

Can the learner solve a similar problem tomorrow?

Can they explain why their solution works?

Can they recognise when an AI-generated answer is wrong?

Can they transfer the concept to an unfamiliar situation?

Do they require less assistance over time?

These may be much more meaningful measures of educational AI.

Instead of optimising only for task completion, we should also be measuring independent capability.


What this means for the AI learning systems I want to build

This question is becoming increasingly important in how I think about ScienceDojo and PracticeDojo.

I don’t want AI to become an answer machine sitting between the learner and the problem.

I want it to behave more like a thoughtful tutor.

That means the system should sometimes explain.

Sometimes question.

Sometimes hint.

Sometimes challenge an assumption.

Sometimes ask the learner to retrieve something they learned earlier.

And sometimes simply wait.

The difficult technical problem is therefore not:

“Can the AI solve this problem?”

Modern AI can already solve an extraordinary number of problems.

The more interesting question is:

“What should the AI do right now to maximise the probability that this learner will eventually be able to solve the problem without it?”

That is a very different design objective.


The goal should be independence

The OECD’s 2026 Digital Education Outlook makes an important argument for the direction of educational AI: generative AI can support learning when it is designed and used with clear pedagogical purpose.

But convenience should not be confused with learning.

A student producing a better answer is useful.

A student developing a better mind is the real objective.

Perhaps one of the best measures of an AI tutor will eventually be surprisingly simple:

As the student learns, they should need the AI less.

And if we can design systems that achieve that, AI may not replace the productive struggle involved in learning.

It may become remarkably good at knowing exactly how much struggle to preserve.


This article was inspired by and discusses findings from the OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education, alongside recent research on productive struggle and pedagogically aligned generative AI. The interpretations and product-design reflections presented here are my own.


Comments

Popular posts from this blog

Why I Chose Automatic Question Generation: The Story Before the Thesis

ScienceDojo Began Before There Was a Website