AI can do the work. But what does that actually prove?

In an earlier article, I wrote about a question that becomes increasingly important as AI can act more independently:

If AI can do something, what may it decide for itself?

While continuing to build and test TiepMiep, the typing course I developed with AI’s help, a second question emerged:

If AI has done the work, how do we know whether the result is good enough?

That may sound technical. Ultimately, though, it is about something very human: what do we base our trust on?

Everything seemed to work

At one point, there was a reasonably functional application.

It had lessons, profiles, scores, progress tracking, feedback and rules to determine whether a learner had sufficiently mastered a component. Various tests had also been run during development.

That felt reassuring.

Until I asked myself a simple question:

What had all those tests actually proved?

They showed that particular parts of the software did what we expected in particular situations.

But the ultimate purpose of TiepMiep was not to have working buttons, scores and lessons.

It was for a child to learn touch typing.

And the software tests had not proved that at all.

Three questions that look similar

As I wrapped up the project, I found it helpful to separate three types of question.

1. Verification — does the system do what we agreed?

Does the functionality work as intended? Is a result saved? Does a learner move to the next lesson after successfully completing an exercise?

2. Process confirmation — can the intended user actually work with it?

Can a child independently select a profile, start a lesson, understand the instructions and complete the exercise?

3. Outcome — do we ultimately achieve what we wanted?

After the course, can the child actually touch-type an unfamiliar Dutch text independently and consistently?

Those are three different claims. Each requires different evidence.

A passing test is not general proof of quality.

AI makes evidence cheap

Something interesting happens here.

With AI, it is becoming easier to design tests, write test code, carry out checks and produce documentation.

That is useful. But it also creates a new risk.

We can easily produce lots of evidence without first being clear about what we are trying to prove.

More evidence is not automatically better evidence.

A folder containing a hundred passing tests may look impressive. But if those tests do not address the main risks or how the system is actually used, the confidence they provide may be smaller than the volume of documentation suggests.

Do not start with the test

This has changed the order in which I think.

Instead of starting with:

Which tests can we run?

I start with:

What do we need sufficient confidence about?

A few simple questions follow:

  • What are we trying to achieve?
  • What could go wrong that matters?
  • How much confidence do we need?
  • What evidence fits that need?

Only then does testing become interesting.

Do not test simply because you can. Test because there is a relevant uncertainty you want to reduce.

And what if something changes?

The same applies to regression testing.

During software development, it is tempting simply to rerun all existing tests after every change. AI makes that increasingly easy technically.

But a thinking step belongs before that too.

What has changed? Which existing functionality or previous evidence might be affected?

That then leads to targeted retesting.

change → impact → affected evidence → targeted retesting

Less spectacular than running hundreds of automated tests. But much clearer about why a test is needed.

The role of Quality Assurance is changing

For me, this connects directly to Quality Assurance.

As AI can take over more execution work, the human contribution shifts.

Beyond checking whether activities were performed, we need to ask:

  • Which claim are we trying to support?
  • Which risk are we trying to control?
  • What evidence do we need?
  • Which conclusion can that evidence actually support?

AI can write code, run tests, analyse results and prepare documentation.

But a large pile of correctly performed activities does not yet amount to a sound conclusion.

What could I ultimately say about TiepMiep?

At the end of this development phase, I could say with reasonable support that there was a functionally implemented prototype, and that a substantial part of its functionality had been verified within the situations tested.

There were also a few observations of children working with it.

But I could not say it had been proved that the course teaches children to touch-type unfamiliar Dutch text independently and consistently.

The evidence simply was not there.

Perhaps that is the most important lesson of the whole experiment:

Quality means more than knowing what works. It also means knowing what you know, what you do not yet know, and which conclusion the available evidence can actually support.

From autonomy to assurance

My previous article ended with:

If AI can do something, what may it do independently?

This experiment adds a second question for me:

If AI has done something, what evidence do we need before we dare trust it?

The first question concerns autonomy.

The second concerns assurance.

Perhaps this is becoming an increasingly important human role in a world where AI can perform more work:

Determining what matters, what evidence fits, and which conclusion that evidence can actually support.