SERVICE / ENGINEERING REVIEW

Liquid-Cooling Design Review

The brief in motion

The Liquid-Cooling Design Review, Explained

Macro view of liquid cooling cold plates and manifold tubing on dark server hardware with amber coolant glow
The cooling loop is infrastructure, not an accessory.

An honest air ceiling beats a flawed liquid loop

Liquid cooling is no longer optional at AI densities, but a flawed loop design is worse than admitting the air ceiling and stopping there. Undersized coolant distribution units, optimistic flow assumptions, single points of failure hidden inside manifolds, and commissioning plans that skip failure-mode testing all look fine on a drawing. They surface later as change orders, thermal throttling, or a flooded row. The review exists to catch them while they are still lines on paper.

Why the crossover is not negotiable

Disciplined air containment holds to roughly 30-40kW per rack. A modern AI training rack draws 80-100kW, and a fully populated NVIDIA GB200 NVL72 rack draws 120-132kW. Past the low end of that range, air stops being an efficiency question and becomes a ceiling no amount of fan power moves past. Every design this review examines exists because the rack density already left air behind.

What the review examines

An independent engineering review of a direct-to-chip or immersion cooling design covers loop architecture, including topology, manifold design, isolation points, and serviceability; CDU sizing validated against real thermal loads rather than nameplate figures; flow rates, pressure drops, and coolant chemistry assumptions; redundancy and failure modes, meaning what actually happens when a pump, a CDU, or a facility loop drops out; the commissioning plan, including leak testing, flow balancing, and load ramp; and serviceability, meaning how a technician swaps a cold plate at two in the morning without draining a row.

Direct-to-chip or immersion: different tradeoffs, not a default answer

Direct-to-chip cold plates target the highest-heat components while preserving a standard server layout and standard serviceability. Immersion cooling submerges entire assemblies in dielectric fluid, maximizing thermal transfer and eliminating fans entirely, at the cost of a fundamentally different service model. Neither is automatically correct. The review states which one your design actually needs and why, rather than assuming the vendor's default.

Measured against what has survived in the field

The review is grounded in two decades of building and shipping liquid-cooled GPU systems. As Technical Director of EKWB USA, Jon Moen designed and delivered $500K liquid-cooled GPU servers to customers including MIT, Cornell, CrowdStrike, the US Navy, and NATO, environments where a cooling failure is not an inconvenience. Your design is measured against what has actually survived in the field, not against what a datasheet claims.

What you receive, and where it starts

A written review that names design defects specifically and ranks them by consequence, CDU and loop sizing validated or corrected against your actual thermal loads, and a commissioning checklist your integrator signs against. Reviews begin from $9,500 as a one-time engagement, scoped after an intake call and a first pass over the design package. It fits naturally after a readiness assessment and before contracts are signed, which is exactly when an independent set of eyes is cheapest. Many clients pair it with Deployment Oversight, so the engineer who validated the design also witnesses its commissioning.

Next step

Get a second set of eyes before capital moves

Send the current design package, and Jon will scope the review from there.