Yesterday, my LLM tool made a small mistake.
I had shared a photo of a handwritten reflection. The tool misread one word and then used that misreading in the reflection it wrote for me.
I corrected it.
The tool looked at the photo again and confidently offered a different interpretation.
Wrong again.
Nothing particularly dramatic. My handwriting is not always easy to read, and it was only one word.
But what happened next was interesting too.
The problem was not that my LLM tool could not produce a good story.
The problem was precisely that it could.
The reflection was logical. The connections sounded plausible. The text flowed nicely. Had I not known better, there would have been little reason to question it.
Only part of the story was built on something that was not right.
Plausible output is not the same as reliable output.
The mistake came before the answer
With AI, we often look at the final product.
Is the answer correct? Does the code work? Is the summary good? Did the test produce the right result?
Reasonable questions.
But something else happens before every answer.
Something is observed. Meaning is assigned to it. Relevance is determined. Possibilities are weighed up. A choice is made.
Only then do we see the answer.
With my handwritten reflection, things went wrong at the very first step: the tool had not correctly observed what was written.
What happened afterwards was actually quite good.
And that is precisely what makes it interesting.
A system that perceives something incorrectly and then produces nonsense is easily caught out.
A system that perceives something incorrectly and then produces a convincing story is much harder to spot.
Do not fill the gap immediately
The strange thing is that I have been encountering the same idea in quite different places lately.
When delivering training, for example.
As a trainer, it is tempting to fill a silence. You know the material. You can see where the conversation could go. Often, you already have an answer ready.
But by saying nothing for a moment, something else can happen.
Someone asks a question you had not considered. A participant puts it differently. Interaction emerges, rather than just knowledge transfer.
The interesting part is not the answer.
It is the space before it.
I am also trying to approach my own reflections more consciously. First, look at what is actually happening. Do not immediately turn every observation into a conclusion, action or improvement point.
Some things can be left for a while.
And then AI started building software
The same question returned in an experiment in which I started developing software with AI.
At first, that is mainly fascinating.
You describe what you want. AI helps develop requirements. Code appears. Functionality is added. Tests are run. Documentation emerges.
The obvious question is:
Can AI do this?
Increasingly, the answer is: quite a lot.
But as building became easier, other questions began to interest me more.
Why are we running this test?
Who decided what was important enough to test?
What does a passing test actually prove?
And if AI writes the code, designs a test, runs it and then concludes that the result is sufficient — where did independent judgement take place?
Gradually, my question shifted.
Not just: what can AI take over?
But also: what is AI actually taking over?
Where was the human judgement?
A workflow consists of more than actions.
It also contains moments when someone observes, interprets, doubts, weighs things up and decides.
Who determined what mattered?
Who noticed that something unusual might be relevant?
Who decided that the available evidence was sufficient?
Who said: I do not really know enough about this yet?
When we automate such a workflow, we may therefore be automating more than work.
We may quietly automate a piece of human judgement too.
That need not be wrong. If a machine can demonstrably perform something better, faster and more consistently, we do not need to insert a human on principle.
But automating an action is different from handing over the judgement hidden within it.
That is why two questions have started to diverge for me:
Can AI perform this step reliably?
and:
Do we want to hand the judgement in this step over to AI?
They look similar.
But they are entirely different questions.
A human at the button
We like to solve this with a human in the loop.
A human remains involved. Problem solved.
But suppose AI gathers the information, interprets what is relevant, weighs up possibilities, formulates a conclusion and presents a proposal.
Then a human sees:
Approve?
Formally, there is a human in the loop.
But how much human judgement have we actually retained?
Perhaps we should look less at whether there is a human somewhere in the process, and more at where human thinking and decision-making actually happen.
It is a subtle difference.
But as AI becomes more capable, I think it becomes increasingly important.
When answers become cheap
This brings me to something that may be bigger still.
For a long time, much knowledge work has been organised around the ability to produce good answers.
The trainer knows what to explain.
The consultant knows what makes sense.
The programmer knows how something should be built.
The quality professional knows what to look for.
Knowledge and experience help you reach a good answer faster.
But AI changes something.
Producing answers is becoming cheaper.
Not flawless. Not risk-free. And certainly not without human oversight.
But fast, convincing and available on an enormous scale.
Perhaps that also shifts what makes human expertise valuable.
Not just knowing the answer.
But looking carefully before answering.
Noticing that something is wrong.
Distinguishing what you actually know from what you make of it.
Understanding which evidence supports a conclusion — and which does not.
Not immediately filling a silence.
Recognising doubt as useful information.
And perhaps above all:
knowing when you have not yet seen enough to be entitled to give an answer.
Perhaps that is where the craft lies
I now encounter this idea in a surprising number of places.
In training.
In quality management.
While building software with AI.
And even during a personal reflection at the end of an ordinary day.
Technology can increasingly produce the answer.
Perhaps that makes the moment before the answer more interesting.
The moment when you look.
Ask another question.
Leave something open.
Doubt.
Or decide that the available evidence does not yet tell you enough.
Perhaps part of what we call human judgement lies right there.
For now, that leaves me with one question:
What happens to human judgement when the answer is no longer the difficult part?
I do not yet have a definitive answer. But I will certainly keep exploring it.