The impact of IT/OT convergence becomes most visible inside the operational workflows teams rely on every day. In the final blog of our IT/OT convergence series, we examine the workflows that shape reliability most directly in AI/HPC environments.
Much of the conversation around IT/OT convergence focuses on system integration: connecting platforms, centralizing telemetry, and improving visibility across infrastructure and compute environments. Those capabilities matter. But the operational impact of convergence becomes tangible inside the workflows teams execute daily — and whether those workflows can hold up when infrastructure systems, IT systems, and operational teams are required to function as a single coordinated model.
That challenge is particularly acute in AI/HPC environments, where infrastructure behavior and compute behavior influence each other in real time. A cooling issue may begin affecting workload performance before infrastructure alarms escalate. Power instability may surface first through application behavior rather than inside the electrical system itself. Operators need the ability to correlate signals across systems quickly enough to understand what’s happening before conditions deteriorate.
Across the industry, the realization is taking hold that IT/OT convergence is actually an operational coordination initiative, and three workflows sit at the center of it.
Workflow 1: Incident Response
Incident response is where the strengths and weaknesses of IT/OT convergence become most visible most quickly. When you think about traditional environments, infrastructure incidents and IT incidents are typically managed as separate events. AI/HPC environments have reduced that separation considerably. A thermal imbalance in a liquid cooling system may begin affecting GPU performance before temperatures cross alarm thresholds; a power quality issue may surface first through workload instability or unexpected compute behavior before facilities teams identify the underlying condition.
Effective incident response in converged environments increasingly depends on shared operational visibility: infrastructure telemetry, workload behavior, operational history, and maintenance activity evaluated together rather than in parallel by separate teams.
Consider a pattern we’ve seen play out in live AI clusters:
During a planned utility transfer test, a UPS static switch oscillated between its inverter and bypass paths for roughly 74 seconds. From the IT side, nothing looked wrong; compute load never dropped, and the power team logged a clean, successful test. But the brief undervoltage transients, though well within IT tolerances, were enough to trip a coolant distribution unit (CDU) pump bank feeding two rows at peak H100 training load. Chilled-liquid flow fell, GPU inlet temperatures climbed from 22°C to over 28°C in eleven minutes, and the racks entered sustained thermal throttling, ultimately failing a checkpoint cycle on a job that had been running for hours.
The instructive part isn’t the transfer test, the pump, or the GPUs. It’s that three teams were each looking at their own system of record:
- The power team saw a successful transfer
- The cooling team saw a pump trip
- The compute team saw throttling GPUs
Each signal lived on a different platform and a different clock (power quality captured in milliseconds, building controls in seconds, GPU telemetry every fifteen minutes) and no single workflow stitched them into one causal chain. The time lost reconstructing “what actually happened, and in what order” was time the training run couldn’t recover. The incident was cross-domain, but the response was single-domain.
The coordination challenge compounds at portfolio scale. Different facilities may escalate incidents differently, classify events differently, and document activity using inconsistent workflows. Over time, leadership loses the ability to evaluate operational performance consistently across sites because the operational record itself becomes fragmented.
Workflow 2: Maintenance and Change Management
Maintenance and change management become significantly more complex in converged environments because infrastructure systems and compute systems are interdependent in ways they weren’t before.
A maintenance activity that once affected only cooling infrastructure may now directly influence workload performance. Firmware updates, cooling adjustments, and electrical maintenance can all introduce downstream effects across interconnected systems (effects that may not be immediately visible inside the system where the work was performed).
Consider a scenario where a cooling control adjustment is made during scheduled maintenance. In a traditional environment, success is measured by whether the cooling system performed as expected. In a converged AI environment, the right question is whether workload temperatures stabilized as expected, whether the adjustment introduced any unexpected compute behavior, and whether downstream thermal conditions held within operating tolerances. The validation requirement extends across both domains simultaneously.
This is where operational discipline becomes the differentiator. Teams need workflows capable of validating not only whether maintenance was completed, but whether the intended operational outcome was achieved, across both infrastructure and compute, afterward.
The organizations making the most progress here are connecting procedures, operational records, telemetry, and post-maintenance validation into a coordinated workflow rather than treating each activity as a discrete, domain-specific event.
Workflow 3: Asset Visibility and Performance Management
Asset visibility becomes substantially more valuable when operational data from multiple systems is evaluated together.
Telemetry provides continuous visibility into asset behavior: temperatures, power utilization, cooling performance, alarms, thresholds, environmental conditions. Maintenance records provide operational context: what work was performed, when conditions changed, how assets have behaved historically. Independently, each dataset provides part of the picture. Together, they create something more useful: an operational understanding of how infrastructure is actually performing over time, not just right now.
This matters most in AI/HPC environments, where small performance deviations can begin affecting workload behavior long before infrastructure systems identify a formal failure condition. A recurring cooling anomaly that appears insignificant when viewed as isolated alarm activity may, when correlated against maintenance history and workload performance data, reveal early signs of degradation that would otherwise remain invisible until they become something harder to recover from.
Take a coolant distribution unit (CDU) as an example:
In isolation, a CDU that occasionally reports a supply temperature a degree or two above setpoint, or a pump that momentarily works a few percent harder to hold flow, throws alarms minor enough to be acknowledged and cleared without a second thought. Viewed as isolated alarm activity, all it is is noise, but when that same CDU’s telemetry is read against its maintenance history (a coolant top-up here, a filter change there, a heat-exchanger approach temperature that has been quietly widening across successive PM cycles) and against the workload it’s cooling, the pattern reads very differently. It’s early-stage degradation of a pump or heat exchanger, visible weeks before it would surface as the flow-loss failover that throttles a GPU row. Whether that gets caught as a maintenance-planning item or as a 2 a.m. incident comes down to one thing: whether telemetry and operational history are evaluated together.
This is precisely the shift the most advanced operators are making. At one hyperscale portfolio, a global command center fields on the order of seven million alarms a year, with only a fraction of a percent tied to genuine SLA risk. No team triages that by reacting to thresholds. The advantage is not more alarms; it is correlating the few that matter against maintenance history and workload behavior, so a slow-developing asset problem is surfaced as a trend long before it becomes a failure.
The operational advantage IT/OT convergence creates here is the ability to move from isolated visibility toward coordinated intelligence, and for portfolios expanding across facilities, providers, and increasingly dense compute environments, that shift is what separates proactive management from reactive response.
Building Toward a More Integrated Operating Model
As AI data center environments continue evolving, the operational relationship between infrastructure systems and compute systems will only continue to tighten. One in five significant outages now costs operators more than $1 million, and that number is rising as AI/HPC workloads increase both the financial exposure tied to downtime and the speed at which infrastructure conditions can affect revenue-generating compute performance.
The organizations making the most progress on IT/OT convergence are treating it as a workflow and operational coordination challenge, rather than a system integration. They’re building operational models where telemetry, maintenance activity, procedures, escalation workflows, and asset history contribute to a shared understanding of how the environment is functioning in real time, across both IT and infrastructure domains.
The conversation doesn’t end here. This summer, we’re unpacking more operational challenges impacting mission-critical teams every week. Visit our insights page to read our latest blogs and subscribe for monthly updates.