Enterprise Platforms
Reliability is not a platform property
September 2026
A service can be technically healthy and still be operationally unreliable. The difference is often found in what happens around the technology.
The situation
The service supported a globally distributed workforce using real-time collaboration, telephony, video, meeting-room and enterprise communication services.
Delivering those services meant operating across a chain of dependencies that no individual engineering team completely controlled.
A failed call might involve the collaboration platform, an endpoint, a local network, regional connectivity, authentication, telephony infrastructure or a service provider. A poor meeting experience could originate in the room, the device, the network or the platform itself.
Each component could appear healthy in isolation. The employee could still have a bad experience. From their perspective, the distinction was irrelevant. The service either worked or it did not.
The operating problem
Traditional operational models make it relatively easy to assign ownership of technology.
They are less effective at assigning ownership of outcomes.
When reliability is organised around components, each team can investigate its part of the stack and conclude that it is operating normally. An incident can therefore move between teams even though nobody has yet explained the failure experienced by the user.
The organisational boundaries become part of the technical problem. That led to a different question.
How do you operate reliability across a system whose failure modes do not respect organisational, supplier or geographic boundaries?
What changed
The approach shifted from looking primarily at individual platform health towards understanding the end-to-end service.
That meant correlating evidence across different parts of the environment rather than relying on a single platform's telemetry to explain what users were experiencing.
Recurring incidents also needed to be treated as patterns rather than isolated tickets.
If apparently unrelated failures repeatedly appeared in particular locations, network paths, devices or service dependencies, that pattern was itself evidence. The objective was to identify where reliability was actually being lost rather than simply determining which team should receive the next ticket.
This also changed the nature of operational ownership.
Teams still needed clear technical responsibilities, but resolving an end-to-end problem sometimes required somebody to remain accountable for the outcome after their own component had been ruled out.
Result / evidence
The immediate benefit was better diagnosis.
Engineers could distinguish more quickly between failures originating in the collaboration platform and those caused elsewhere in the service chain. Recurring patterns became easier to identify, and investigations could begin with evidence rather than assumptions about which technology was responsible.
The larger improvement, however, was conceptual.
Platform availability and service reliability stopped being treated as interchangeable measures.
A system could satisfy its component-level operational metrics and still fail the person trying to use it.
That distinction changed how incidents were investigated and how service health was understood.
What I took from it
Reliability is not an inherent property of a platform. It is an emergent property of the system around it.
A cloud service can meet its availability target while an employee still cannot make a call. A meeting-room device can report online while the room remains unusable. A network can satisfy its infrastructure metrics while application performance is poor.
This is why mature service management cannot stop at component availability. The useful unit of analysis is the experience produced by the whole system.
Sometimes the highest-value reliability improvement is not another platform change. It is better telemetry, clearer ownership, stronger dependency mapping, better coordination between suppliers or a faster way to correlate evidence across technical domains.
The difficult part is that none of those things necessarily belongs to the platform team. The service still depends on them.
The operating question
If every component says it is healthy but the user still cannot work, who owns the failure?