Most colocation operators have invested heavily in their monitoring infrastructure. The gap that’s actually costing them is harder to see, and almost never where they’re looking.
According to the Uptime Institute’s 2025 Annual Outage Analysis, nearly 40% of organizations have suffered a major outage caused by human error in the past three years, and 85% of those incidents traced back to procedures not being followed or, arguably more troubling, not having adequate procedures in the first place.
Walk into almost any modern colocation data center and you’ll find thousands of data points watched in real time, and power, cooling, humidity, thresholds, alarms all visible in real time, doing exactly what it was designed to do.
So if monitoring isn’t the problem, and the BMS is working as it should, what is the problem? You’ll find it in the seven minutes after the BMS fires:
- Who gets the alarm?
- How does it travel from the monitoring system to the person who needs to act on it?
- Is there a work order, or a phone call?
- Is the technician who responds following a documented procedure or working from memory?
- And when the event is resolved, whether it took seven minutes or seven hours, what record exists that can withstand a tenant inquiry, compliance review, or post-incident debrief with your VP?
For directors managing colocation portfolios at scale, this blind spot between the monitoring layer and the operational layer is one of the most consequential and least examined sources of risk in the business. It accumulates one underdocumented event at a time, until something happens that makes it visible.
With outages now carrying financial exposure that can exceed $200,000 per hour in AI/HPC environments (and one in five significant outages now costing operators more than $1 million), the operational blind spots between the alarm and the response are becoming significantly harder for operators to absorb.
The Blind Spot in Action
The BMS-to-CMMS blind spot isn’t a technology problem in the traditional sense, because colocation facilities typically have both systems. The gap is in the handoff between them, and that handoff is almost entirely dependent on human behavior and workflow discipline.
In practice, it manifests in three ways:
- Latency: the time between a BMS alarm firing and a technician receiving a contextualized work order is still measured in hours rather than minutes in many environments.
- Context loss: technicians arrive with the alert itself but no fault history, trend data, or prior maintenance context, making resolution more dependent on individual experience than operational history.
- Documentation drift: across portfolios with multiple FM providers and teams, the same event gets handled and documented differently between sites, making it difficult for leadership to know whether performance is actually improving or just being recorded inconsistently.
The following anonymized incident illustrates how quickly the gap between alarm and action becomes visible (and how it often takes a tenant escalation to create that visibility).
Anonymized Incident Report
Incident Type: Cooling alarm response delay / BMS-to-CMMS handoff gap
Environment: Multi-tenant colocation facility, high-density tenant room
Severity: Near miss; no sustained outage; tenant-visible thermal degradation
Primary Issue: Critical alarm acknowledged in BMS but not converted into an actionable, traceable work order
Timeline:
02:13 BMS generates critical cooling alarm
02:15 Alarm acknowledged in BMS
02:18 Technician contacted verbally; no CMMS work order created
02:27 Technician arrives with alarm label but no asset history or procedure
02:34 Tenant monitoring detects elevated inlet temperatures and opens urgent ticket
02:41 Manual work order created after tenant escalation
02:48 Cooling unit issue isolated; adjacent cooling capacity verified
03:05 Temperature trend stabilizes
Post-event: Incident timeline reconstructed manually from BMS logs, phone notes, tenant ticket, and technician updates
Root Cause:
The operational failure was the disconnected handoff between BMS alarm, CMMS work record, asset history, and field execution. The alarm existed. The work did not become traceable until after the tenant escalated.
Contributing Factors:
- No automatic or established manual work order creation from a critical BMS alarm
- Prior fault history not visible to the responding technician
- Asset naming mismatch between BMS and CMMS
- Initial response depended on verbal handoff instead of a documented workflow
The Alarm Volume Problem Nobody Talks About
One dimension contributing to this blind spot that deserves more attention than it gets is sheer volume. A large colocation portfolio may generate tens of thousands of alarms daily across BMS, EPMS, and supporting monitoring systems. Over time, teams inevitably begin filtering alarm volume informally just to keep pace.
The problem is that those unofficial triage habits rarely exist in documented workflows, which leaves leadership with little visibility into how response prioritization is actually happening across the environment. Even when the BMS is working perfectly, the operational layer has adapted around it in ways that are invisible to leadership… until they aren’t.
What the Tenant Sees That You Don’t
In a colocation environment, the BMS-to-CMMS gap has a dimension that doesn’t exist in enterprise facilities: tenants are watching too.
Sophisticated tenants, like those coming from hyperscaler environments or running AI and HPC workloads, have their own monitoring infrastructure: they can see performance degradation before an alarm fires, and identify thermal anomalies, power fluctuations, and availability metrics in real time. What they can’t see is whether the event was caught, documented, escalated correctly, and resolved in a way that prevents recurrence.
The moment a tenant identifies an event through their own monitoring before hearing from you, the dynamic of that relationship changes. The question quickly shifts from “did you resolve it?” to “why did we know before you did?”
That gap (between the event and the communication, between the resolution and the documentation, between what happened and what you can prove happened) is entirely an operational layer problem. The BMS did its job. What follows is a reflection of the maturity of the systems, workflows, and disciplines that sit between the alarm and the tenant conversation.
The Questions Worth Sitting With
Directors running colocation portfolios rarely lack confidence in their monitoring infrastructure. The questions that are harder to answer with confidence are the ones that sit on the other side of the BMS:
- When an alarm fires at Site 7 at 3 AM, how long does it actually take for a documented work order to exist? Is the answer the same at Site 12? At Site 23?
- When a tenant asks for an incident history on a specific asset, how long does it take to produce one, and how complete is it?
- When a near-miss is resolved at one site, does that resolution inform how a similar event is handled at the next one?
The monitoring infrastructure passed the test a long time ago; these questions are about what comes after. If you want a clearer picture of how your operation performs under that kind of scrutiny, learn how MCIM’s Operational Audits can help.