Systems
The Pilot Is Not the Product
13 min
A pilot demonstrates possibility. A product must survive reality.
The Pilot Is Not the Product
There is a point in almost every successful technology pilot when the conversation changes. The system works. Users have tried it. The feedback is positive. Perhaps there is a number showing how much time has been saved or how much faster a process has become. What began as an experiment starts to look like something the organisation should scale.
That is usually the right moment to ask a slightly different question.
What exactly did the pilot prove?
A pilot is designed to reduce uncertainty. Can this technology do what we think it can? Is the use case useful? Will people actually use it? Is there enough value here to justify further investment? Organisations need those experiments because committing fully before answering those questions would be expensive and often irresponsible.
The problem is that success in a pilot can begin to mean more than the evidence supports. A pilot shows that something can work under the conditions created for the pilot. A product has to keep working when many of those conditions are no longer there.
That difference sounds obvious. Inside organisations, it is surprisingly easy to lose.
The conditions around the technology
Pilots usually take place in relatively controlled environments. The number of users is limited. The use case is narrow enough to understand. The people involved often know they are testing something new, which changes how they respond when it behaves unexpectedly. The project team is close to the technology and problems tend to receive attention quickly.
None of this is a criticism of pilots. It is what makes them useful.
Experiments become difficult to learn from if every possible organisational complication is introduced at once. But those conditions matter when we later interpret the result.
Imagine an internal AI assistant is tested by fifty employees. It answers questions using company information and, over several weeks, users report that they are finding answers faster. The project team records good adoption and positive feedback. Perhaps the assistant saves each user twenty minutes a day.
That tells us something important. There may be real value in the capability.
It does not yet tell us what happens when five thousand people use it, when the original project team is no longer watching every issue, when documents change without anybody remembering that they feed an AI system, or when employees begin asking questions that were never represented in the original test.
The underlying technology may be exactly the same. The system around it is not.
This distinction has existed in engineering for a long time. Google's Site Reliability Engineering practice, for example, uses Production Readiness Reviews when assessing whether a service is ready to be operated reliably. Those reviews consider things such as monitoring, dependencies, emergency response, capacity and change management. The fact that a service functions is not, by itself, enough for an SRE team to take responsibility for running it.
There is a simple idea underneath that practice.
Working technology and an operable service are not the same thing.
Some of the pilot may actually be people
There is another reason a pilot can give us more confidence than it should.
People compensate for weaknesses in systems all the time.
An engineer notices an integration failure and fixes it before most users see it. A project manager chases an access request that has become stuck. Someone checks a dataset manually before loading it. A subject-matter expert knows that one source contains outdated information and makes sure nobody relies on it. A developer investigates an unusual result immediately because there are only a handful of users and the problem has appeared in a project chat.
Individually, those interventions may seem insignificant.
Together, they can form part of the operating model without anybody recognising them as such.
This is particularly easy to miss during a pilot because the people around the technology are often unusually knowledgeable and motivated. They know what the system is supposed to do. They know its weaknesses. They know who to call when something fails. Often they helped build it.
The pilot therefore contains more capability than appears on the architecture diagram.
Some of that capability lives in people.
Once the pilot succeeds, those people may move on. The project closes. A supplier completes its engagement. Engineers return to other priorities. The technology remains, but some of the mechanisms that made it reliable quietly disappear.
At that point the organisation discovers whether those mechanisms were temporary support or whether they were actually part of the product.
I think this is one of the more useful questions to ask before scaling anything:
What are people currently doing manually, informally or exceptionally to make this work?
The answer does not mean everything needs to be automated. Some exceptions are better handled by people. Some activities may already exist elsewhere in the organisation. Others may be so rare that building a permanent capability around them would make no sense.
But the work has to be visible before somebody can make that decision.
Otherwise it simply vanishes at the end of the pilot.
Scale changes the system
We often talk about scale as though it were multiplication.
Something worked for one hundred people, so now the challenge is to make the same thing work for ten thousand.
But ten thousand people do not behave like one hundred people repeated one hundred times.
Larger populations introduce variation. People interpret instructions differently. They discover uses that designers never anticipated. They encounter combinations of circumstances that did not appear in a small test. Some people trust technology too much. Others distrust it and create workarounds. Different teams begin incorporating it into processes the project never knew existed.
The technology can also change behaviour simply by being useful.
Suppose an AI tool really does save employees twenty minutes every day. Initially, that is a productivity improvement. But people adjust to improvements. Processes are redesigned around the faster way of working. Expectations change. What previously took an hour may gradually be expected to take forty minutes.
Eventually the organisation is no longer receiving twenty minutes of optional benefit.
It is assuming the twenty minutes will always be there.
The capability has become a dependency.
There is rarely a meeting where somebody formally declares that this transition has happened. It tends to emerge gradually. A tool becomes convenient, then normal, then expected. Only when it stops working does the extent of the dependency become obvious.
That is why successful adoption changes the risk profile of a technology.
The more valuable something becomes, the more consequential its failure can become too.
AI creates a different operating problem
This becomes more complicated with generative AI because failure is not always easy to recognise.
Traditional services have given operations teams a reasonably familiar set of questions. Is the application available? Is it responding quickly enough? Are requests failing? Are the infrastructure and dependencies healthy? Is capacity sufficient?
Those questions still matter for AI systems.
They are no longer enough.
An AI service can be technically healthy and still be performing badly. Authentication works. The API responds. Latency is normal. The database is available. No alerts are firing. Every conventional dashboard can be green while the answers users receive are becoming less useful.
Microsoft's current guidance on generative-AI observability reflects this difference. It includes normal operational measures such as latency and errors, but also evaluation of output quality using measures such as groundedness, relevance, safety and task completion. Microsoft explicitly describes production monitoring as something that must track both operational health and quality in real-world use.
That changes what operating a service means.
We are accustomed to monitoring whether software is functioning.
With AI, we increasingly need some way of understanding whether the behaviour produced by that functioning software remains acceptable.
That is much harder.
Consider a simple internal assistant answering questions from company documentation. An employee asks about a policy and receives the wrong answer.
What failed?
Perhaps the model generated an incorrect response. Perhaps the correct document was never available to the retrieval system. Perhaps an old document remained indexed after a newer version was published. Perhaps the source itself contained incorrect information. Perhaps the system retrieved the right evidence but failed to use it properly. Perhaps a permission change meant the relevant information was no longer accessible.
Or perhaps the answer is not technically wrong at all. The user may have asked a question the service was never designed to answer.
"AI gave the wrong answer" describes the symptom.
It tells us very little about the failure.
A pilot can tolerate that ambiguity because the people who built the system are nearby and there may only be a few incidents to investigate. At scale, repeated bespoke investigation becomes expensive very quickly.
This is why observability for AI cannot only mean more logs. The organisation needs some way of understanding the path between the question, the information available to the system, the behaviour of the model and the answer eventually shown to the user.
Otherwise the organisation knows that something went wrong but not where to look.
What did we actually test?
The conditions of an experiment also influence what its results mean.
NIST's AI Risk Management Framework recommends evaluating systems in conditions similar to those in which they will be deployed and maintaining post-deployment monitoring once those systems are in use. It also includes incident response, recovery, change management, user feedback and decommissioning as part of managing AI systems after deployment.
That matters because a pilot can produce a very convincing number while leaving important assumptions hidden.
Imagine an AI assistant achieves 95% accuracy during testing.
It sounds excellent.
But who wrote the test questions?
Were they created by people who already understood how the system worked? Were difficult or ambiguous requests represented? Did the dataset contain the same inconsistencies as the information the system will encounter in production? Were users trying to accomplish real work, or were they following test scenarios designed by the project?
Perhaps 95% is genuinely impressive.
The point is that the number cannot tell us by itself.
The meaning of the result depends on the conditions that produced it.
The same problem appears outside AI. A service tested in one office tells us relatively little about how it will behave across hundreds of locations with different connectivity. A new process tested entirely by experienced employees may hide problems that become obvious when somebody joins the organisation. A support model that works while five engineers know every component personally may behave very differently once responsibility is distributed across several teams.
The more production differs from the test environment, the more carefully the evidence from the test needs to be interpreted.
Production eventually becomes an ownership problem
One of the clearest differences between a pilot and a product appears when something goes wrong.
During a pilot, the project normally owns the problem.
In production, "send it to the project team" eventually stops being an answer.
Somebody has to own the service. Somebody has to understand its dependencies. Somebody has to know what information is needed to investigate a problem. Somebody has to decide whether a change is safe. Somebody has to know what happens when the supplier changes something, when the model is updated, or when the source information feeding the system changes.
This is not simply a matter of allocating a support queue.
Ownership carries knowledge.
If a service desk receives a ticket saying that an AI assistant is giving incorrect answers, what can they actually do with it? Can the application team see what documents were retrieved? Can somebody determine which version of the information was used? Can the business owner establish what the answer should have been? Is there a known escalation route when the issue is with the underlying model rather than the application?
If those questions have no answer, the product may technically be in production while the organisation around it is still operating like a project.
That distinction becomes particularly visible months later, when the people who remember all the original decisions are no longer involved.
A successful pilot should create more questions
There is a strange asymmetry in how organisations respond to pilots.
When a pilot fails, people become curious.
Was the use case wrong? Was the technology unsuitable? Did people struggle to use it? Was the data poor? Were our assumptions wrong?
Failure creates questions.
Success can sometimes stop them.
Once something has been labelled successful, concerns about ownership, monitoring or support can begin to sound like obstacles to progress. We proved the value. Why are we slowing this down?
But a successful pilot should not end the questioning.
It should change the questions.
Which conditions made the pilot work?
Which of those conditions will still exist at scale?
Which ones depended on temporary people or processes?
What new behaviours appear when far more people use the system?
How will we know when performance is gradually deteriorating rather than simply unavailable?
Who will notice when one of the assumptions behind the original design is no longer true?
And perhaps most importantly, what happens if the technology is so successful that the organisation starts depending on it?
These are not arguments for endless governance.
They are questions about what has to become real when an experiment becomes infrastructure.
So what?
A pilot should not have to solve every possible production problem. That would make experimentation unnecessarily slow and expensive.
The purpose of a pilot is to learn enough to make the next decision.
But that next decision should distinguish between evidence of value and evidence of readiness.
Before scaling something that appears successful, I would want to understand at least a few things. What conditions made the pilot successful? What manual work is currently being performed around the technology? What becomes different when the number and variety of users increases? How will the organisation recognise degradation rather than complete failure? Who owns the service after the people who built the pilot move on? And what happens if adoption creates a dependency that did not previously exist?
Sometimes the answers will be reassuring.
An existing operational model may already cover most of what is needed. A service may genuinely be simple enough that moving from pilot to production introduces little additional complexity.
Sometimes the questions will reveal work that still needs to be done.
That is not evidence that the pilot failed.
It may be exactly what a successful pilot was supposed to reveal.
What begins after the pilot
Pilots create visible progress. There is a launch, a group of users, a demonstration, results and usually a decision.
Production is less visible.
Much of the work happens precisely so that nobody has to think about it. Monitoring detects a problem before users do. A support team knows where to send an incident. Documentation means an engineer does not have to find the person who originally built the system. A dependency is changed without breaking something downstream. An old source is removed before anybody receives the wrong answer.
When these things work, they can look like overhead.
They are not.
They are part of the product.
The interface is part of the product. The model is part of the product. The automation is part of the product.
So are ownership, telemetry, support, documentation, controls, recovery and the ability to understand what happened when something goes wrong.
A pilot asks whether an idea deserves to continue.
A product has to survive the consequences of continuing.
The real test begins when the conditions that made the pilot successful are no longer exceptional.
Perhaps that is the question worth asking when the successful demonstration ends.
Not simply whether the technology can scale.
Whether everything around it can scale too.
Sources and further reading
[1] Cruz, A. and Bhambhani, A. (2016) “The Evolving SRE Engagement Model,” in Site Reliability Engineering. Google.
Describes Google's Production Readiness Review model and the operational areas considered before SRE assumes responsibility for a production service, including monitoring, dependencies, emergency response, capacity and change management.
[2] Microsoft (2026) “Observability in Generative AI,” Microsoft Foundry documentation.
Describes observability across the AI lifecycle, including conventional operational telemetry alongside evaluation of groundedness, relevance, safety and task completion in production.
[3] Microsoft (2026) “Observability for Generative AI and agentic AI systems,” Microsoft Learn.
Discusses why conventional observability is insufficient for probabilistic AI systems and why evaluation and AI-native signals need to complement traditional logs, metrics and traces.
[4] National Institute of Standards and Technology (2023) Artificial Intelligence Risk Management Framework (AI RMF 1.0).
Includes post-deployment monitoring, feedback, incident response, recovery, change management and decommissioning as part of ongoing AI risk management.