I keep hearing the same phrase: we must keep a human in the loop.
And I understand why.
When an AI is making a decision that could affect somebody's job, health, money or future, we do not want an unaccountable machine wandering off on its own.
But I have started to think that the word human is doing far too much work in that sentence.
A human can be tired. A human can be distracted. A human can misunderstand the system, work outside their expertise, accept the recommendation because challenging it takes too long, or approve something they do not have enough information to judge.
Humans get things wrong.
I have personal reason to know that. Six years ago, a doctor failed to carry out a test that should have been done, and the consequences nearly killed me. The fact that a human professional was involved did not make the decision safe.
So I think we need a better question.
Is there demonstrated competence in the loop for this exact decision?
Being human is a biological fact. Being competent is an evidence-based claim.
Human presence is not the same as human oversight
A person appearing at the end of a process and clicking approve may satisfy a workflow diagram. It does not prove that they have understood what happened.
The UK's Information Commissioner's Office already makes this distinction in data-protection guidance that is particularly relevant to automated decisions with legal or similarly significant effects. Meaningful human involvement should not be a token gesture. The reviewer should actively weigh the recommendation, consider the available information, and have the authority and competence to go against it. The ICO marks that guidance as under review, and it should not be treated as a universal legal definition for every AI workflow.
The EU AI Act follows the same direction for high-risk systems. As appropriate and proportionate to the circumstances, its human-oversight provisions require people to understand capabilities and limitations, remain alert to automation bias, interpret outputs, disregard or override them where appropriate, and stop the system.
That is a much stronger idea than merely saying, "a human was involved".
It says the human must be able to do something useful.
Putting a human and an AI together does not guarantee a better answer
This matters because human plus AI does not automatically equal superhuman.
A 2024 systematic review and meta-analysis examined 106 experiments reporting 370 effect sizes from studies published between January 2020 and June 2023. Each experiment compared human-only, AI-only and combined performance. On average, human-AI combinations performed better than humans working alone, but worse than whichever of the human or AI performed best alone.
The average result was augmentation, but not synergy.
The study also found substantial variation. The type of task mattered. Decision tasks showed average performance losses from the combination, while creation tasks looked more promising, although the estimated gain for creation tasks was not statistically different from zero in that subset.
That does not mean collaboration is a bad idea.
It means we must test the configuration. We cannot assume that adding a person improves the system, or that adding an AI improves the person.
Competence is narrow
I can be competent in one area and completely useless in another.
A brilliant accountant is not therefore a brilliant fitness trainer. A capable engineer is not therefore a capable dentist. A good programmer may know very little about fixing an engine. A chief executive's authority does not create technical expertise.
So "reviewed by an expert" is not precise enough.
What was the decision? What evidence did it require? What could go wrong? What must the reviewer be able to notice, challenge and stop?
Competence has to be attached to a defined activity under defined conditions.
That is why a degree, job title, certificate or number of years in a role is useful evidence, but not complete proof. Medicine recognises this. A medical degree is not the end of professional formation; supervised practice, appraisal, continuing development and revalidation continue because capability has to remain current and relevant to the work being done.
I think AI governance needs the same humility.
Competence is necessary. It is not sufficient.
Here is where the phrase needs another turn.
A highly competent person can still be a useless safeguard if they do not have the evidence, time, independence or power to act.
For oversight to be real, I think five things have to be present together:
- Competence: can this person or team assess this exact class of decision?
- Information: can they see the evidence, uncertainty, limitations and relevant history?
- Capacity: do they have enough time and attention to investigate rather than rubber-stamp?
- Authority: can they override, pause, stop or escalate the system without being punished for it?
- Accountability: is someone named for the outcome, with an appeal and learning route when it goes wrong?
I would treat these as an AND test, not a score that lets strength in one area cancel out absence in another.
A famous expert with no time is not effective oversight. A diligent junior person with no authority is not effective oversight. A board member with accountability but no domain knowledge is not effective oversight.
The danger of the moral crumple zone
There is another reason to be careful about putting a person into the diagram.
Researcher Madeleine Clare Elish uses the phrase moral crumple zone for situations where the human closest to an automated system absorbs the blame, even though they had limited practical control over its design, information or behaviour.
The 2018 fatal collision involving an Uber automated test vehicle is a sobering example of why the system around the person matters. The US National Transportation Safety Board found that the operator failed to monitor the road while distracted by her phone. It also identified organisational contributors: inadequate safety-risk assessment, ineffective oversight, insufficient safeguards against automation complacency and an inadequate safety culture.
The lesson is not that the human did not matter.
It is that naming one human at the end of a badly designed system does not repair the system.
A confidence tree for one decision
The second question is how we prove competence and keep proving it.
I have been exploring something I call a confidence tree. I do not mean a machine-learning decision tree, and I am not claiming this is an established regulatory standard. It is a practical assurance structure for making the evidence behind a decision visible.
At its root is a bounded claim:
We are justified in trusting this human-AI configuration to make or oversee this decision, in these conditions, using this system version, until this review date.
The branches then test the claim:
- Has task performance been demonstrated on realistic cases?
- Does the reviewer understand this AI's capabilities, failure modes and uncertainty?
- Can they identify when the case is outside their competence?
- Can they intervene in time and exercise real stop authority?
- Are decisions, overrides, complaints and outcomes recorded and reviewed?
- What would invalidate the evidence, and when must it be reassessed?
This fits with the NIST AI Risk Management Framework, which calls for human-AI roles to be defined, operator and practitioner proficiency to be assessed and documented, relevant limitations to be available to decision-makers, and performance to be evaluated in conditions similar to deployment.
The important word is demonstrated.
We should not ask whether somebody attended the training. We should ask whether they can detect a plausible error, explain the system's limits, make a sound decision under pressure, know when to escalate, and use the stop mechanism correctly.
And then we should check again later.
The most competent entity may be a configuration
Once we frame the problem this way, the answer does not always have to be one heroic expert.
The most competent decision-maker might be:
- a human working alone;
- an AI carrying out a narrow, well-tested task;
- a domain expert assisted by AI;
- an AI supervised by a person with specific stop authority;
- two systems checking different failure modes;
- or a multidisciplinary team where no single member holds all the required competence.
The choice should depend on measured performance, risk, law, legitimacy and who is affected. In consequential settings, human and institutional accountability cannot be outsourced to a model merely because the model performs well.
But nor should we insert a person into every loop and assume the problem has been solved.
The questions I would ask a board
When somebody presents an AI process with a box marked human approval, I would ask:
- What exact decision is this person being trusted to oversee?
- How have you demonstrated their competence for that decision?
- What do they know about this model and its failure modes?
- What evidence and uncertainty can they actually see?
- How much time do they have per case?
- Can they disagree, override and stop the process?
- What happens when they are uncertain or unavailable?
- Who is accountable, and how can the affected person appeal?
- How do you measure whether the combined human-AI system is better?
- When does that evidence expire?
If we cannot answer those questions, we do not have competent oversight.
We have a human-shaped decoration in the loop.
So yes, keep people involved. But make the claim honest.
I am not arguing for removing people from consequential decisions.
I am arguing for taking their role seriously.
The objective is not maximum human involvement. It is the best achievable decision quality, safety, legitimacy and accountability.
Sometimes that will require more human judgement, not less. Sometimes it will require a narrower specialist. Sometimes it will require a team. Sometimes the AI should do the routine work while a human handles exceptions. Sometimes the system should stop because nobody currently has enough evidence to proceed.
Human oversight is not automatically a safety mechanism.
Competent oversight can be.
Related reading
- You Are The A In AI
- When Do We Need Judgement?
- The AI Screwed Up. Then What?
- How Long Will Agentic Work Take?
This is a proposed governance principle and practical test, not legal, medical or professional advice. The correct oversight and accountability requirements depend on the decision, jurisdiction and sector.
