What AI governance can learn from Quality Assurance

And why we should avoid repeating all the pharmaceutical industry’s mistakes.

AI systems are getting more scope to act independently. They search for information, combine sources, advise people and can use tools to take real actions.

This raises an important question:

How do we know such a system operates under sufficient control?

AI companies talk about alignment, guardrails, evaluations, monitoring, human oversight and AI governance.

When I read about this recently, something felt familiar.

Many of these problems seem strikingly familiar to me.

After all, in the pharmaceutical industry, we have spent decades trying to answer a similar question:

How do we create sufficient confidence in a complex system without trusting it blindly?

Perhaps AI governance does not need to reinvent everything.

But a warning is appropriate too. If AI governance learns from Quality Assurance, let it adopt our best lessons — rather than our thickest SOPs.

When AI does something it was not supposed to do

OpenAI recently published a framework for reporting model misalignment incidents: situations where a model displays unexpected or unwanted behaviour.

The examples are fascinating.

In one case, during an assignment, a model found and used a publicly exposed API key without authorisation. When it still could not find the requested information, it fabricated data.

In another, an agent had found the correct answer with Python but still needed an online source to cite. Its solution: independently upload the file to the internet.

OpenAI describes a process for identifying, investigating, classifying and disclosing such events. The company itself emphasises that the published cases are individual incidents and do not indicate how often this behaviour occurs.

Reading it, I immediately thought of my own field.

In Silicon Valley, we call it AI governance. In pharma, we would probably just open a deviation.

That is an exaggeration, of course. But the underlying idea is interesting.

Something unforeseen happens. You want to know what actually occurred. What was the impact? Can we reconstruct what the system did? Was this isolated or systemic? Do we need to intervene immediately? What can we learn? Are additional measures needed?

Suddenly, that sounds surprisingly familiar.

Deviation management. Root cause analysis. CAPA. Change control. Monitoring.

AI governance meets Quality Assurance.

Suppose we deploy an AI agent

Let us make this concrete.

Suppose an organisation introduces an AI agent to help employees assess quality incidents.

Initially, the agent does little. It gathers information from various documents and produces a summary.

Later, it gets more capabilities.

It classifies the incident, finds similar cases, proposes possible causes and advises on follow-up action.

Later still, it gains access to systems and may enter information or prepare actions.

The same agent can therefore take on very different roles:

summarise → advise → prepare → act.

And the risk changes with them.

An incorrect summary checked by an experienced employee is very different from an agent that can independently modify a quality record.

Yet we often ask an overly general question about AI:

“Is this AI safe?”

From a Quality Assurance perspective, I would rather ask:

What is the intended use, what could go wrong, and how much confidence do we need that the relevant risks are sufficiently controlled?

From testing to assurance

With computerised systems, we have long tended to create trust mainly through extensive testing and documentation.

Computer Software Assurance, or CSA, represents an interesting development in that thinking.

The FDA uses CSA specifically for software used in medical-device production and Quality Management Systems. It is therefore not simply a general pharmaceutical rule for every AI application.

The underlying principle is interesting, however: use a risk-based approach to establish sufficient confidence in software and apply greater rigour where the risk justifies it.

That sounds logical.

Practice is more stubborn.

Even with a risk-based approach, a familiar reflex easily emerges:

“This component is actually low risk, but let us test it fully anyway, just to be safe.”

Better safe than sorry.

Before you know it, you have formally introduced a risk-based method while still testing almost everything in practice.

I see the same danger in AI governance.

More testing does not automatically mean more assurance.

Large volumes of test cases, approvals and documentation may even distract attention from the risks that really matter.

The better question is therefore not:

How much should we test?

But:

Which uncertainty are we trying to remove with this test?

Stop saying AI has been “tested”

This brings us to a problem we know well from software validation.

We use the word testing for all kinds of activities.

Functional testing. Regression testing. User Acceptance Testing. Performance testing. Security testing.

AI adds model evaluations, robustness testing, adversarial testing and red teaming.

But these activities answer different questions.

Take our AI agent.

A functional test can demonstrate that it retrieves information from the right system.

A regression test can investigate whether a change has degraded existing functionality or previously demonstrated performance.

An AI evaluation can measure how reliably the agent classifies quality incidents within a relevant dataset.

A control test can demonstrate that the agent cannot close a quality record without the required authorisation.

And during User Acceptance Testing, we can examine whether users can adequately carry out the intended workflow with the overall system.

All tests.

But all provide different evidence.

This may seem a detail, but it is essential.

Because when someone says:

“The AI has been extensively tested.”

We still know almost nothing.

The interesting follow-up is:

Which assurance question did those tests answer?

We may already have some of the necessary evidence too.

Supplier tests, automated regression tests, earlier evaluations and operational data can all provide relevant information.

The skill is not rerunning every test.

It is determining:

What evidence do we need, what reliable evidence already exists, and what relevant uncertainty remains?

Risk-based work has a blind spot too

There is an uncomfortable assumption in risk-based work: that we understand the most important risks sufficiently in advance.

With AI, that need not always be the case.

An agent can encounter a new situation, combine different sources or use a tool in a way the developer had not anticipated.

If we only test what we identified as critical beforehand, we may miss unexpected failure modes.

Risk-based assurance therefore does not mean:

Test only the risks you already know.

It also means actively looking for risks you do not yet know.

Exploratory testing, adversarial testing, red teaming and monitoring play an important role here.

And sometimes extensive testing is simply sensible. If the potential impact is high, uncertainty is considerable and additional testing is relatively easy, it can be an excellent risk control.

The problem is therefore not lots of testing.

It is testing without knowing which uncertainty or risk you are addressing.

And then there is the human

A common solution to AI risk is a human in the loop.

Let AI advise, and let a human make the final decision.

Problem solved.

Or is it?

Suppose our AI agent almost always gives excellent advice.

Initially, the QA employee checks each recommendation carefully.

After a hundred good recommendations, behaviour changes. The employee reads faster. After a thousand, checking may mainly become confirming.

Formally, our control still exists:

AI → human review → approval.

In practice, something else may have emerged:

AI → rubber stamp → approval.

We have designed a human control without demonstrating that it remains effective in practice.

That too is Quality Assurance.

Asking whether a control actually manages the risk it was designed for, as well as whether the control exists.

Validation is not an endpoint

This is where AI becomes interesting.

A generative AI system may be probabilistic. Context changes. Models are updated. Data changes. Agents gain new tools. Users discover new ways to work with them.

That makes the model:

test → approve → done

increasingly difficult to sustain.

Perhaps we therefore need to think more strongly in terms of lifecycle assurance for AI:

understand → assess risk → test selectively → use under control → observe → learn → adapt.

I see an interesting parallel with Continued Process Verification in pharmaceutical process validation.

An AI system is not the same as a production process.

The similarity lies in the principle.

In Continued Process Verification, information is collected and evaluated during commercial production to assess whether the process remains in a state of control.

For AI, similar thinking would mean that sufficient confidence does not come solely from evidence collected before deployment.

We also need to keep learning from actual use.

The question is therefore more than:

Did we have sufficient evidence at deployment?

It also becomes:

Do we still have sufficient evidence?

Monitoring, logging, trends, overrides, abnormal performance and incidents thus become part of assurance.

But not every AI error is a deviation

Here too, we need to avoid simply copying our existing QA reflexes.

Generative AI has variability and can make mistakes.

If we treat every incorrect or unexpected output as a formal deviation, we will probably soon create an unworkable system.

We need to learn to distinguish, for example:

expected variability → abnormal performance → failed control → actual incident.

Deviation management can therefore be a useful concept.

That does not mean every hallucination needs a deviation form.

Perhaps this is the real lesson

AI governance can learn a lot from Quality Assurance.

But perhaps even more from the development Quality Assurance itself has undergone.

Not from compliance to assurance, as though compliance no longer mattered.

Compliance remains necessary. Laws and regulations, procedures, responsibilities and documented controls form the foundation.

But following those rules is not yet proof that a system actually operates under control in practice.

The development I mean is closer to:

compliance as an endpoint → compliance as the foundation for assurance

From:

demonstrating that we followed the procedure → also demonstrating that the procedure actually controls the intended risk

From:

documenting everything → collecting relevant, reliable evidence

From:

extensive testing → understanding which risks and uncertainties our tests are intended to address

From:

validation as a project → assurance throughout the lifecycle

The distinction matters.

A system can be validated entirely according to procedure, contain every required signature and still provide insufficient assurance about actual use.

Conversely, excellent technical evidence can never justify ignoring applicable regulations or necessary governance.

We need both.

Compliance tells us which requirements and agreements we must meet.

Assurance gives us evidence-based confidence that the relevant risks are actually controlled.

Trust needs a basis

This creates an interesting opportunity for AI governance.

We need not repeat the same learning curve to discover why, in computerised systems, we increasingly try to think in terms of risk and assurance.

But we must also be willing to accept the difficult consequences.

Risk-based work sometimes means deliberately not testing something extensively.

Human oversight means more than adding an approval step: it means demonstrating that the human control is effective.

And governance means being able to explain why these particular controls are needed, rather than adding as many as possible.

AI governance therefore does not need to reinvent the wheel.

Quality Assurance has decades of experience with intended use, risk, evidence, suppliers, changes, deviations, monitoring and responsibility.

But Quality has also shown how easily compliance without sufficient attention to assurance leads to false confidence: the procedure was followed, the boxes ticked and the documents approved — while the most important question remains insufficiently answered.

I would therefore not advise AI governance simply to adopt our methods.

Learn from our development instead — including the mistakes we made along the way.

Ultimately, assurance centres on a simple but difficult question:

Do we have sufficient reason to trust that this system, for this application, under these circumstances, is used under sufficient control?

Compliance is not its opposite.

It is the foundation. Assurance then asks whether that foundation holds up in practice.

Perhaps there is a simple way to assess whether we really answer that question:

If you remove all test protocols, ticked boxes and approvals tomorrow — what evidence remains that your AI is actually under control?


Sources

OpenAI — Our framework for reporting model misalignment, 16 September 2026.

U.S. FDA — Computer Software Assurance for Production and Quality Management System Software, February 2026.

U.S. FDA — Process Validation: General Principles and Practices, January 2011.