The GPU hall is not the factory. The factory is the plant that keeps those GPUs inside their thermal envelope. When a SuperPOD trips, the first screen people open is the cluster. The cause is usually earlier: a CDU that lost flow, a primary circuit that had already warmed, or a chiller that was sitting on the edge of its capacity.
You cannot react your way through an NVIDIA AI factory. The heat is dense, the liquid loops have little forgiveness, and the window between "still training" and "throttled or down" is measured in minutes. A CRAH hall used to give you hours. This does not.
Act. Do not wait for the GPU to tell you.
The factory is a cooling chain, not a rack row
An NVIDIA AI factory is a heat machine first. The GPUs, NVLink domains and InfiniBand fabric take the attention because that is where the capital sits. The cooling chain is what keeps that capital in service.
Heat leaves the cold plate into a secondary coolant loop. That loop runs to a Coolant Distribution Unit. The CDU hands the heat across a plate exchanger into the building's primary water. The primary circuit then has to reject it — chillers, dry coolers, pumps, valves, the energy centre.
If you only watch the GPU, you are watching the last thing to fail.
Commissioning still matters. Fabric, power and silicon-speed tests prove the cluster can run. They do not prove it can run at 03:00 when a strainer starts to block and secondary supply temperature is already climbing. That is an operations problem, and it needs a live picture of the plant, not a handover pack.
The CDU is the pinch point
The CDU is where landlord water meets tenant coolant. It is also where most monitoring stops being useful, because the two sides are still on different systems, in different units, with different owners.
On the primary circuit you need supply and return temperature, flow, system pressure and energy valve state. On the secondary circuit you need supply and return, flow per row, differential pressure and pump status — including whether the N+2 set has already taken a pump out. Inside the CDU you need plate exchanger delta-T, expansion vessel level, filter differential pressure and the failover logic actually firing, not the schematic that says it should.
Those points tell you where the heat is stuck.
A rising secondary supply with a falling primary delta-T is not a rack problem. The exchanger is not moving heat, or the primary side is not taking it. If the only number you have is rack inlet, you find out when the GPU derates. That is reacting.
PODVIEW treats the CDU as an object, not a row of analogue values on a BMS graphic. Primary and secondary sit on the same panel. You can see a pump fail over without waiting for the row to warm. You can see filter differential climb while the loop still looks "in spec".

Watch the plant that feeds the CDU
The CDU is only as good as what arrives at its primary connections.
If building water is warm, short on flow, or unstable on pressure, the CDU will pass that problem straight into the secondary loop. The GPU hall then looks like the failure. It is not. The failure started in the energy centre, or in the pipework between the energy centre and the hall.
That means the same board has to show the infrastructure feeding the CDU: chiller and dry cooler state, primary pumps, strainers, isolation valves, energy valves, and the temperatures and flows on the risers — not only the CDU itself. A healthy CDU graphic with falling primary flow is a plant problem you still have a few minutes to catch. A GPU thermal event is the same problem, ten minutes later, with a training job dead.
Demarcation makes this worse if you let it. Landlord owns the primary. Tenant owns the secondary. The CDU sits on the line. If each party only watches their half, nobody sees the hand-off. Clear monitoring boundaries are not a contract footnote. They are the difference between "we acted on primary flow" and "we argued about whose alarm it was while the hall came off."
Act, do not react
Reacting is waiting for a thermal trip, a derate, or a red tile on the cluster dashboard. Acting is seeing the primary circuit move before the secondary does, and the secondary before the rack does.
That only works if the points land in one place, with the same units and the same timestamps. BMS, PDUs, CDUs, leak, power — not four vendor consoles and a spreadsheet. PODVIEW is built for that: points arrive from whatever is already installed, get normalised, and sit on one hall view. CDUs are colour-coded by state. Click one and the circuits open. Walk the 3D floor without walking the floor.
Alerts have to match the chain, not the GPU. Warning on CDU pressure approaching capacity. Warning on plate delta-T collapsing. Warning on primary flow dropping while secondary load holds. Critical when secondary supply is leaving the envelope. By the time the GPU reports a thermal event you are already late.
The same record is what you take to the morning meeting and to the customer. Not "the hall got hot." The primary return was up, the exchanger delta-T was down, pump 2 had already failed over, and the action was taken on the plant — or it was not, and you can see the minutes you lost.
What this does not replace
Watching the cooling chain does not replace NVIDIA's commissioning tests, a competent BMS, or a protection system on the electrical side. It does not command the plant. It tells operations what the plant is doing to the factory, in time to intervene.
It also does not turn a liquid hall into something you can ignore. High-density AI floors have less thermal mass and less time. The monitoring has to be as tight as the cooling design, or the design is a brochure.
PODVIEW already watches CDU primary and secondary circuits, flow, pressure, plate exchanger delta-T and pump failover, alongside the rest of the hall. That is the picture an AI factory needs: not only the cluster, and not only the CDU, but the infrastructure that feeds the CDU that then feeds the factory.
If you only look at the GPUs, you will always be reacting. The factory runs on the water.
