Governance
Evidence Is an Architecture
11 min
What organisations choose to measure becomes the reality they are able to govern.
Evidence is often treated as something that simply exists and waits to be collected. A dashboard reports performance. A control produces evidence. A metric tells us whether something is improving. An audit establishes whether a requirement has been met.
Inside organisations, it is rarely that neutral.
Before anything appears on a dashboard, someone has already made a series of decisions. What should be measured? What counts as success? What threshold matters? Which population is included? How often should the measure be taken? Who sees it? What happens when it turns red?
Those decisions create a field of attention.
Things inside that field become easier to see, discuss, escalate and fund. Things outside it do not necessarily stop existing. They simply become harder for the institution to notice.
That is why I have started thinking of evidence less as an output and more as an architecture.
Measurement does not simply observe
There is a comfortable assumption behind many management systems. First reality happens, then we measure it.
Sometimes that is approximately true. But measurement also changes the environment in which people operate.
Wendy Espeland and Michael Sauder studied what happened when US law schools became increasingly subject to public rankings. Their research describes what they call reactivity: people and organisations change their behaviour when they know they are being measured.[1]
The rankings did not simply describe law schools. They began influencing how schools allocated resources, interpreted success and made decisions.
That mechanism is not difficult to recognise elsewhere.
If a service team is primarily judged on the speed with which incidents are closed, closure speed becomes important. If a customer operation is judged on call duration, call duration attracts attention. If a programme reports the percentage of milestones delivered on time, milestones become part of how success is understood.
None of those measures is necessarily wrong.
The problem begins when we forget that the measure is a representation of the thing we care about rather than the thing itself.
A closed incident is not necessarily a resolved problem.
A short call is not necessarily a satisfied customer.
A project delivered on schedule is not necessarily a useful project.
Numbers create clarity partly by removing context. That is both their strength and their danger.
What can be counted becomes easier to govern
Espeland and Mitchell Stevens describe quantification as a social process, not merely a mathematical one. Turning something into numbers allows comparison across things that may previously have been difficult to compare. It can simplify complexity, create categories and make decisions possible at scale.[2]
Organisations need this.
Nobody running a large service can personally investigate every incident. An executive cannot read every customer complaint before deciding where to invest. Regulators cannot inspect every transaction manually.
We create indicators because institutions need compression. The question is what is lost during that compression. Suppose an organisation reports service availability of 99.95%.
That number is useful. But it cannot tell us by itself whether the 0.05% of unavailable time happened at 3 a.m. on Sunday or during the organisation's busiest trading period. It cannot tell us whether ten people were affected or twenty thousand. It cannot tell us whether the same failure has happened repeatedly or whether the underlying system is becoming progressively harder to operate.
The number is not false. It is incomplete by design. Every measurement system is. The difficulty comes when the compressed representation begins travelling further through the organisation than the context from which it was produced.
By the time information reaches senior governance, a complicated reality may have become a percentage, a red-amber-green status or a single sentence on a slide. That compression is necessary. But we should not mistake it for neutrality.
Measures become targets
There is another problem once evidence becomes connected to consequences.
Donald Campbell was writing about this in the 1970s. His argument, now commonly described as Campbell's Law, was that the more heavily a quantitative indicator is used for social decision-making, the greater the pressure for the underlying process to become distorted around that indicator.[3]
The mechanism is straightforward. If people know what counts, they adapt. Sometimes that is exactly what we want.
A security team begins measuring patching performance and patching improves. A service desk starts tracking abandoned calls and staffing decisions change. A programme introduces clear delivery measures and overdue work receives attention.
Measurement has helped direct behaviour. But optimisation can continue past the point we intended. Teams learn which tickets affect the measure. Definitions get interpreted in ways that improve the number. Exceptions become classifications.
Work that matters but does not appear in the metric becomes harder to prioritise.
Eventually a strange thing can happen: the evidence improves while the underlying condition does not improve at the same rate.
The dashboard becomes greener.
The organisation feels safer.
But part of what has improved is the organisation's ability to produce a green dashboard.
This does not require dishonesty.
People respond rationally to the systems around them.
If an organisation repeatedly tells people that one number represents success, it should not be surprised when activity begins organising itself around that number.
Evidence can become ceremonial
Michael Power's work on what he called the audit society is useful here. He examined the enormous growth of auditing, inspection, quality assurance and other forms of formal verification, and the organisational consequences of making more activities auditable.[4]
There is an important distinction between something being well controlled and something being easy to demonstrate as controlled.
The two often overlap.
They are not the same.
Anyone who has worked inside a heavily governed organisation will have seen some version of this.
A control has an owner.
An assessment happens every quarter.
Evidence is uploaded.
A reviewer signs it off.
The process is complete.
But what does that evidence actually tell us?
Sometimes a great deal.
Sometimes it proves mainly that the process capable of producing the evidence is functioning.
This is where governance can become strangely self-referential. The institution builds a mechanism for verifying control, then starts treating successful completion of the verification mechanism as evidence that the underlying risk itself is understood. That is not an argument against controls or assurance. Organisations need both. It is an argument for asking one question further.
What does this evidence allow us to know that we could not know before?
If the answer is unclear, we may be measuring the governance process rather than the condition the governance process was created to manage.
Absence of evidence is particularly dangerous
There is another failure mode that interests me.
Once an organisation becomes accustomed to governing through evidence, things without evidence can begin to feel as though they do not exist.
No incidents have been reported.
No control failures have been identified.
No customer complaints have crossed the threshold.
No alerts have fired.
That sounds reassuring.
But each statement contains an assumption about the system that produces the evidence.
Would we have seen the problem if it occurred?
Would somebody know how to report it?
Does the monitoring cover the relevant failure mode?
Is the threshold sensitive enough?
Are people incentivised to surface bad news?
Can the system distinguish a genuinely healthy condition from a condition it does not know how to observe?
Monitoring tells us something only if we understand what the monitoring is capable of detecting. A silent alarm is very different from an alarm that detected nothing. Yet dashboards can make those conditions look remarkably similar.
AI makes the architecture of evidence more important
This becomes even more consequential as organisations introduce AI into decision-making. AI systems generate enormous amounts of apparent evidence: evaluations, confidence measures, model outputs, retrieval traces, test results and automated assessments. But more evidence does not automatically create more assurance.
The National Institute of Standards and Technology's Generative AI Profile explicitly recommends documenting data origin and content lineage, evaluating data and content flows, recording dependencies on upstream data sources, and assessing AI outputs against appropriate ground truth and evaluation methods.[5]
That matters because the question is no longer simply whether an AI system produced an answer. We increasingly need to know where the answer came from.
Which sources influenced it?
What transformation occurred between the source and the output?
What was tested?
Under what conditions?
What uncertainty remains?
What would cause us not to trust the result?
These are architectural questions.
An AI system may produce a beautifully documented answer while relying on weak evidence. A retrieval system may cite a document without that document actually supporting the claim being made. A model may perform well against a benchmark that poorly represents the environment in which it will actually operate. The presence of evidence is not the same thing as evidence being meaningful. As AI systems become more involved in operational and organisational decisions, I think this distinction will become much harder to ignore.
The evidence chain matters
Perhaps the most useful way to think about evidence is backwards.
Start with the decision.
What decision are we actually trying to make?
Then ask what we would need to know to make that decision well.
What evidence could tell us that?
How would that evidence be produced?
What could distort it?
What would remain invisible?
Who has an incentive to influence the measure?
What uncertainty needs to travel with the result?
That is very different from starting with whatever data happens to be available and building a dashboard around it.
A dashboard should not exist because a system can produce twelve metrics.
A control should not collect evidence simply because the evidence is easy to capture.
A governance committee should not receive a measure unless the relationship between that measure and the decision it is supposed to inform is understood.
Otherwise information accumulates without necessarily increasing understanding.
Evidence should carry its limitations
One thing organisations could do better is preserve uncertainty instead of removing it as information moves upward. We tend to reward clean numbers. An exact percentage feels more authoritative than an estimate. A single status feels more decisive than several competing interpretations. A green indicator is easier to absorb than a paragraph describing uncertainty. But sometimes uncertainty is part of the evidence.
Perhaps the sample is small.
Perhaps the monitoring has a blind spot.
Perhaps two data sources disagree.
Perhaps the definition changed halfway through the reporting period.
Perhaps the metric measures a proxy because the actual outcome is difficult to observe.
Removing those qualifications may make information easier to consume while making the resulting decision worse. Good evidence does not merely tell us what we know. It helps us understand how well we know it.
So what?
If evidence is an architecture, then organisations should design it with the same care they apply to other important systems.
Start with decisions rather than dashboards. Use more than one signal when the thing being governed is complex. Distinguish an outcome from a proxy for that outcome. Keep enough context to understand what a number hides. Ask how people are likely to behave once a measure becomes consequential. Test whether evidence-producing controls actually reveal the condition they are supposed to govern. Preserve provenance so that important claims can be traced back to their source. And periodically ask what the organisation cannot currently see. That last question may be the most important. Because every evidence architecture has blind spots. A mature governance system is not one that pretends otherwise. It is one that knows where some of them are.
What becomes governable
Evidence is necessary because institutions cannot operate entirely through direct observation. Complexity has to be compressed. Decisions have to be made. Responsibility has to be demonstrated.
The danger is forgetting that the evidence system has shaped what reached us.
A metric contains choices.
A threshold contains judgement.
A dashboard contains a model of what matters.
An audit trail contains a model of what should be reconstructable.
An AI evaluation contains assumptions about what good performance looks like.
None of this makes evidence less valuable.
It makes evidence something we need to design deliberately.
What an institution measures does not merely describe reality. It helps construct the reality the institution is able to govern.
Perhaps the most important evidence is sometimes not the number in front of us.
It is understanding how that number came to be there, what disappeared along the way, and what decisions it now makes possible.
Sources and further reading
[1] Espeland, W.N. and Sauder, M. (2007) “Rankings and Reactivity: How Public Measures Recreate Social Worlds,” The American journal of sociology, 113(1), pp. 1–40. Available at: https://doi.org/10.1086/517897.
Examines how organisations change behaviour in response to public measures, using US law-school rankings to develop the concept of reactivity.
[2] Espeland, W.N. and Stevens, M.L. (2008) “A Sociology of Quantification,” Archives européennes de sociologie. European journal of sociology., 49(3), pp. 401–436. Available at: https://doi.org/10.1017/S0003975609000150.
Examines quantification as a social process and the ways numerical representations enable comparison, communication and institutional decision-making.
[3] Campbell, D.T. (2011) “Assessing the Impact of Planned Social Change,” Journal of multidisciplinary evaluation, 7(15), pp. 3–43. Available at: https://doi.org/10.56645/jmde.v7i15.297.
Develops the argument commonly known as Campbell's Law: consequential indicators become subject to pressures that can distort both the indicator and the process being monitored.
[4] Power, M. (1999) The audit society : rituals of verification. Oxford ; Tokyo: Oxford University Press. Available at: https://doi.org/10.1093/acprof:oso/9780198296034.001.0001.
Examines the expansion of auditing and verification practices and the organisational consequences of designing activities so that they can be formally inspected and demonstrated.
[5] Autio, C. et al. (2024) Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. DOI: 10.6028/NIST.AI.600-1.
Provides guidance on generative-AI risk management, including data provenance, content lineage, documentation, evaluation, upstream dependencies and testing against appropriate ground truth.