The Operating Question

Decision-making

Judgement Is a System

June 202612 min

Good decisions are not instincts. They are practices, constraints and feedback.

Good decisions are not instincts. They are practices, constraints and feedback.

Organisations often talk about judgement as though it were an individual quality. Someone is experienced. Someone has good instincts. Someone can be trusted to make the difficult call when the process no longer provides an obvious answer.

There is some truth in that. Experience matters. Expertise matters. People do get better at recognising situations they have encountered before.

But judgement is also heavily shaped by the environment in which a decision is made: what information reaches the person, which signals are treated as important, whether disagreement is welcome, how much time is available, what incentives surround the decision and what happens afterwards if the decision turns out to be wrong.

That makes judgement more than a personal capability. It makes it, at least partly, a property of the system.

Experience matters, but experience is not enough

Gary Klein's research on naturalistic decision-making began by studying people making consequential decisions in real environments: firefighters, military commanders and other professionals working with uncertainty, incomplete information and time pressure.

One of the important findings from that tradition is that experienced decision-makers do not always compare a long list of alternatives. They often recognise patterns. A situation looks sufficiently similar to something encountered before, suggesting a plausible course of action that can then be mentally tested against what might happen next.[1]

This is very different from the idea of intuition as something mysterious.

Good intuition is often compressed experience. But that immediately creates another question.

When should we trust it?

Daniel Kahneman and Gary Klein approached intuition from very different research traditions. Kahneman's work emphasised the biases that can distort human judgement. Klein's work studied situations in which experienced professionals made remarkably effective decisions without formal analysis.

When they examined where they actually disagreed, they found considerable common ground.

Their conclusion was that reliable intuitive expertise requires at least two important conditions: an environment with enough regularity to be learnable, and sufficient opportunity for the person to learn those regularities through experience and feedback.[2]

That distinction matters.

Someone can have twenty years of experience in an environment that provides poor feedback and still become confidently wrong.

Experience only becomes expertise when the environment teaches. That is an organisational problem as much as an individual one.

The system decides what the decision-maker gets to see

Consider the person making a decision during a major operational incident. We may later describe the outcome as their judgement.

But what did they actually have available at the time?

Which monitoring had fired?

Which information had reached them?

Were multiple teams looking at different parts of the same failure?

Was there a clear picture of business impact?

Did someone know about a previous incident with similar symptoms?

Were engineers comfortable saying that the prevailing theory did not make sense?

Did escalation bring additional expertise or simply more seniority?

The quality of the decision cannot be separated cleanly from those conditions. This applies outside incidents too.

An executive deciding whether to fund a programme is only as informed as the evidence brought into the room. A risk owner assessing whether a control is effective depends on the quality of the assurance underneath it. A service owner deciding whether to accept a change relies on the assumptions, testing and unresolved risks that have made it through the delivery process.

We tend to locate the final decision in the individual. But the decision was being shaped long before it reached them.

Information architecture is part of judgement.

Governance is part of judgement.

Escalation is part of judgement.

What an organisation allows people to know affects what it allows them to decide.

Challenge is part of the decision system

There is another condition that matters: whether people can disagree.

Amy Edmondson's research on psychological safety defines it as a shared belief that a team is safe for interpersonal risk-taking. In her study of 51 work teams, psychological safety was associated with learning behaviour: asking for help, discussing errors, seeking feedback and engaging in the behaviours through which teams learn.[3]

That matters for judgement because difficult decisions rarely arrive with perfectly complete information.

Somebody may notice something inconsistent.

Somebody may have evidence that does not fit the dominant explanation.

A less senior engineer may understand one technical dependency better than the person chairing the meeting.

A colleague may simply have a bad feeling based on experience and not yet be able to explain it perfectly.

If the organisational environment makes challenging the prevailing view expensive, those signals become less likely to travel.

The decision-maker can still be intelligent, experienced and well intentioned.

They simply never receive the information necessary to exercise good judgement.

This is why "speak up" is an inadequate control.

The useful question is whether the system makes speaking up consequentially safe.

What happens when someone challenges a senior person?

What happens when they delay a delivery?

What happens when they surface a risk everybody hoped had disappeared?

What happens when their concern eventually turns out to be wrong?

If people are punished for false alarms, the organisation should expect fewer alarms.

And eventually it may get exactly what it appears to want: fewer problems being raised.

That is not the same as having fewer problems.

Expertise and authority are not always in the same place

High-reliability organisations offer another useful way of thinking about this.

Research into organisations operating in environments where mistakes can be catastrophic has repeatedly looked at how expertise is distributed and mobilised. Karl Weick and Karlene Roberts' work on aircraft-carrier flight decks described reliable performance not simply as the result of exceptional individuals, but as a form of coordinated attention across interdependent people.[4]

Later work on high-reliability organisations developed the idea of deference to expertise: when circumstances become unusual, decision-making should be able to migrate towards the people with the most relevant knowledge rather than automatically staying with whoever occupies the highest position in the hierarchy.[5]

That sounds obvious.

In practice, it can be difficult.

Hierarchy is persistent. Organisational charts are clear. Expertise is contextual.

The person with the most authority over a service may not be the person who understands the particular failure occurring at 2 a.m. The architect who designed a system may know less about its current behaviour than the engineer who has supported it for three years. A programme director may know exactly what the delivery plan says while an operational team knows which assumption is already beginning to fail.

Good judgement therefore requires more than having experts somewhere in the organisation.

The organisation needs ways of finding them and allowing their expertise to matter at the right moment.

Escalation should not only move decisions upwards.

Sometimes it needs to move them sideways or downwards, towards knowledge.

Feedback turns experience into judgement

There is a reason the feedback condition in Kahneman and Klein's work matters so much. Without feedback, a decision-maker can easily learn the wrong lesson.

Suppose a risky change goes ahead and nothing immediately breaks.

Was the decision good? Perhaps.

Or perhaps the team was lucky.

Perhaps the risk was overstated.

Perhaps the control worked.

Perhaps the failure has a long latency period.

Perhaps somebody quietly intervened and prevented the impact.

Outcome and decision quality are not the same thing. Bad decisions sometimes produce good outcomes. Good decisions sometimes produce bad ones.

If organisations judge decisions purely from outcomes, they make it difficult for people to calibrate their judgement.

This is why structured reflection matters.

Research on after-event reviews has found that experience alone is not enough for learning. Deliberately examining what happened, why it happened and what should be carried forward can improve the value organisations extract from both success and failure.[6]

That requires more than conducting a post-incident review because policy says one is required.

The useful questions are uncomfortable ones.

What did we believe at the time?

What evidence supported that belief?

What did we miss?

Which assumptions turned out to be wrong?

Which signal did somebody notice but fail to escalate?

What worked because the decision was sound, and what worked because we were fortunate?

What would we do differently if the same situation appeared tomorrow?

That is how judgement becomes cumulative. Otherwise an organisation can accumulate years of events without accumulating equivalent wisdom.

Process should support judgement, not replace it

It is tempting to respond to inconsistent judgement by introducing more process.

Sometimes that is exactly the right answer. Checklists, decision criteria, peer review, approval thresholds and escalation routes can protect people from predictable errors. They make expectations explicit and reduce dependence on memory. But there is a point where process starts pretending that uncertainty has disappeared.

A process can tell us which questions should normally be answered. It cannot guarantee that we have asked the right question for an unfamiliar situation.

A risk score can help structure attention. It cannot necessarily tell us which risk is about to become important.

An approval framework can ensure the right functions have been consulted. It cannot make those functions challenge poor assumptions.

The purpose of good process should be to create better conditions for judgement, not to eliminate the need for it.

That means knowing where standardisation is valuable and where people need room to interpret the situation in front of them.

The more complex the environment, the more important that distinction becomes.

AI does not remove the judgement problem

Artificial intelligence makes this especially interesting.

It is easy to imagine that better models will reduce our dependence on human judgement. In some areas they probably will. Systems can analyse more information, detect patterns and produce recommendations faster than people can.

But the introduction of AI also creates new judgement questions.

When should a recommendation be trusted?

When should it be challenged?

Who understands the limitations of the model?

What evidence should accompany the output?

What happens when the human disagrees?

Does the organisation genuinely expect people to exercise oversight, or has the presence of the AI effectively shifted the burden of proof onto anyone willing to contradict it?

NIST's AI Risk Management Framework reflects this problem. It explicitly calls for organisations to define human roles and responsibilities in AI decision-making and oversight rather than assuming that putting a person somewhere in the process automatically creates meaningful human control.[7]

This distinction matters.

A human being present at the end of an automated process is not necessarily exercising judgement. If they lack the information, authority, time or confidence to challenge the system, they may simply be approving its output.

The same principle applies to human and machine decision-making.

Judgement depends on conditions.

So what?

If judgement is partly a system property, then organisations can design for it.

That does not mean creating a committee for every difficult decision. It means looking carefully at the conditions surrounding consequential decisions.

Do the right people receive the right information early enough to use it?

Can someone challenge the prevailing view without paying an unnecessary social or professional cost?

Does authority move towards relevant expertise when circumstances change?

Are assumptions visible enough to be questioned?

Do decisions create feedback that helps people learn whether their reasoning was sound?

Is there enough institutional memory to recognise patterns without blindly repeating previous solutions?

Do processes protect against known failure modes while leaving room for unfamiliar ones?

And when a decision proves wrong, does the organisation ask how the system shaped that decision, or does it simply search for the person who made the call?

That last question matters.

Accountability is necessary.

But accountability that stops at the individual can prevent learning about the conditions that made the error likely.

If several capable people repeatedly make similar mistakes inside the same environment, the interesting question may no longer be why those individuals exercised poor judgement.

It may be what the organisation keeps teaching reasonable people to do.

The conditions for judgement

There will always be moments when organisations need someone to make a call without complete information.

No framework removes that.

No amount of data removes uncertainty.

No AI system removes responsibility.

Good judgement matters precisely because some decisions cannot be completely specified in advance. But that does not mean organisations should simply hope that the right person happens to possess it.

They can build better conditions around the decision.

They can make useful information visible.

They can preserve dissent.

They can connect authority to expertise.

They can create feedback.

They can remember why previous decisions succeeded or failed.

They can give people enough structure to protect against predictable mistakes without creating so much structure that nobody notices when the situation has changed.

Judgement is still exercised by people.

But organisations decide whether those people are operating in an environment that helps them see clearly, challenge assumptions and learn.

That is why judgement is not only something people possess.

It is something systems can strengthen or quietly destroy.

Sources and further reading

[1]Klein, G. (2008) “Naturalistic Decision Making,” Human factors, 50(3), pp. 456–460. Available at: https://doi.org/10.1518/001872008X288385.
Reviews the naturalistic decision-making research tradition and the role of experience and pattern recognition in decision-making under real-world conditions involving uncertainty, time pressure and complex goals.

[2]Kahneman, D. and Klein, G. (2009) “Conditions for Intuitive Expertise: A Failure to Disagree,” The American psychologist, 64(6), pp. 515–526. Available at: https://doi.org/10.1037/a0016755.
Examines when professional intuition is likely to be reliable, emphasising the regularity of the environment and whether decision-makers have sufficient opportunities to learn through feedback.

[3]Edmondson, A. (1999) “Psychological Safety and Learning Behavior in Work Teams,” Administrative science quarterly, 44(2), pp. 350–383. Available at: https://doi.org/10.2307/2666999.
A multimethod field study of 51 work teams examining psychological safety and its relationship with learning behaviours including seeking feedback, discussing errors and asking for assistance.

[4]Weick, K.E. and Roberts, K.H. (1993) “Collective Mind in Organizations: Heedful Interrelating on Flight Decks,” Administrative science quarterly, 38(3), pp. 357–381. Available at: https://doi.org/10.2307/2393372.
Examines reliability on aircraft-carrier flight decks as an organisational accomplishment created through careful, interdependent action rather than simply individual performance.

[5]Weick, K.E. and Sutcliffe, K.M. (2015) Managing the Unexpected: Sustained Performance in a Complex World. Third edition. Edited by K.M. Sutcliffe. Hoboken, N.J: Jossey Bass. Available at: https://doi.org/10.1002/9781119175834.
Develops principles associated with high-reliability organising, including sensitivity to operations, reluctance to simplify, commitment to resilience and deference to relevant expertise.

[6]Ellis, S. and Davidi, I. (2005) “After-Event Reviews: Drawing Lessons From Successful and Failed Experience,” Journal of applied psychology, 90(5), pp. 857–871. Available at: https://doi.org/10.1037/0021-9010.90.5.857.
Examines structured after-event review as a mechanism for learning from experience rather than assuming experience itself automatically produces learning.

[7] National Institute of Standards and Technology (2023) Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. Available at: https://doi.org/10.6028/NIST.AI.100-1.
Includes guidance on defining and differentiating human roles and responsibilities in AI decision-making and oversight, and recognises the organisational and cognitive factors affecting human-AI interaction.

Judgement Is a System | The Operating Question