| IT Heat Load |
Total rack power, server power density, average load, and peak load |
- Conventional air-cooled racks: commonly below 15–20 kW per rack
- High-density liquid-cooled racks: commonly 20–100+ kW per rack
- Design for measured peak demand plus future growth
|
Use liquid cooling when rack heat density exceeds the practical or economical capability of room air cooling. Size the system for peak thermal load rather than average utilization. |
| Cooling Architecture |
Cooling method, heat-transfer path, and level of liquid contact with IT equipment |
- Direct-to-chip cooling: coolant flows through cold plates attached to processors
- Rear-door heat exchangers: liquid-assisted heat removal at the rack exhaust
- Immersion cooling: servers are placed in a dielectric fluid
|
Choose direct-to-chip cooling for targeted high-power components, rear-door heat exchangers for easier retrofit, and immersion cooling when very high density and low fan power are priorities. |
| Cooling Capacity |
Heat removal capacity of cold plates, manifolds, coolant distribution units, and heat rejection equipment |
- Define capacity in kW per server, rack, row, and cooling loop
- Include design margin, commonly 10–20% above calculated peak load
|
Match the cooling chain to the highest expected IT load. Verify that pumps, heat exchangers, piping, and secondary cooling equipment have sufficient capacity at the required flow and temperature. |
| Coolant Supply Temperature |
Temperature delivered to the IT cooling loop and allowable return temperature |
- Many direct-to-chip systems operate with supply water around 18–32°C
- Higher supply temperatures may enable more hours of economizer operation
- Actual limits depend on server and cold-plate specifications
|
Set supply temperature according to the narrowest equipment requirement. Avoid unnecessarily cold coolant, which increases chiller energy use and condensation risk. |
| Flow Rate |
Coolant flow required to transfer the design heat load |
- Required flow depends on heat load, coolant specific heat, and supply-return temperature difference
- For water, a 10°C temperature rise requires approximately 0.024 L/s per kW of heat
|
Calculate flow for each cold plate, server, rack, and distribution branch. Confirm that the lowest-flow branch still meets processor and accelerator cooling requirements. |
| Pressure and Hydraulic Design |
Operating pressure, pressure drop, pump head, balancing, and maximum allowable pressure |
- Use equipment-specific pressure limits for cold plates and quick-disconnect fittings
- Design branches with balanced flow and measurable differential pressure
|
Keep operating pressure below the lowest component rating. Include isolation valves, pressure relief, drain points, air vents, and leak detection at appropriate locations. |
| Coolant Quality |
Fluid composition, conductivity, corrosion control, filtration, and biological control |
- Use treated water or an approved water-glycol mixture as specified by the equipment supplier
- Control particulate contamination and corrosion products
- Monitor pH, conductivity, temperature, and fluid level
|
Do not mix incompatible coolants or elastomer materials. Establish sampling, filtration, flushing, and replacement procedures before commissioning the system. |
| Condensation Control |
Relationship between coolant temperature, dew point, humidity, and exposed surface temperature |
- Maintain coolant and component surface temperatures above the room dew point
- Use dew-point sensors in areas where condensation could occur
|
Provide automatic temperature control, humidity monitoring, insulation where required, and alarms for low coolant temperature or high room dew point. |
| Facility Water Interface |
Connection between the technology cooling loop and the building cooling-water system |
- Use a cooling distribution unit or heat exchanger where separation is required
- Define primary and secondary loop temperatures, pressures, and water chemistry separately
|
Separate facility water from IT coolant when water quality, pressure, or reliability requirements differ. Confirm compatibility with chillers, dry coolers, cooling towers, and economizers. |
| Reliability and Redundancy |
Availability target, redundant pumps, cooling units, power feeds, and bypass paths |
- Common designs include N, N+1, or 2N capacity arrangements
- Provide maintenance isolation without shutting down critical loads where required
|
Base redundancy on the data center availability objective and the consequence of coolant-system failure. Test failover, pump operation, valve positions, and emergency shutdown sequences. |
| Leak Detection and Protection |
Point sensors, rope sensors, drip trays, automatic isolation, and alarm integration |
- Monitor manifolds, hose connections, cold-plate interfaces, CDU areas, and rack bases
- Configure local and remote alarms with documented response actions
|
Install detection beneath or near every potential leak source. Use quick-disconnects with dripless or low-spill designs and provide safe procedures for draining and replacing components. |
| Server Compatibility |
Supported processors, accelerators, memory modules, storage devices, hoses, manifolds, and rack interfaces |
- Confirm allowable coolant temperature, flow, pressure, fluid type, and connection geometry
- Verify whether power supplies, networking equipment, and storage remain air-cooled
|
Obtain mechanical, electrical, thermal, and firmware compatibility data for every server platform. Do not assume that all components in the same rack support the same liquid-cooling conditions. |
| Rack and Floor Layout |
Rack dimensions, manifold location, piping routes, service clearance, floor loading, and access paths |
- Reserve space for distribution units, valves, filters, sensors, and maintenance access
- Verify static and dynamic floor-load limits for high-density racks
|
Coordinate liquid piping with power distribution, cable trays, fire protection, and airflow paths. Keep serviceable components accessible without removing unrelated equipment. |
| Controls and Monitoring |
Temperature, flow, pressure, conductivity, leak status, pump speed, and alarm management |
- Integrate monitoring with the building management system and data center infrastructure management platform
- Record trend data for thermal performance and preventive maintenance
|
Define alarm thresholds, escalation paths, sensor calibration intervals, and automatic responses. Use independent protection for critical leak and over-temperature events. |
| Energy Efficiency |
Pump power, chiller operation, free cooling potential, fan power, and total cooling-system efficiency |
- Liquid cooling can reduce server fan power and support higher supply-water temperatures
- Evaluate total cooling energy using facility-level metrics such as PUE
|
Compare total facility energy, not only the efficiency of the liquid loop. Optimize pump control, temperature setpoints, heat rejection, and economizer operation together. |
| Maintenance Strategy |
Filter replacement, coolant testing, pump service, hose inspection, flushing, and component replacement |
- Define preventive maintenance by operating hours, fluid condition, and manufacturer requirements
- Keep critical seals, hoses, fittings, pumps, sensors, and coolant available as spare parts
|
Document isolation, draining, refilling, purging, and leak-check procedures. Train technicians to work on liquid systems without contaminating IT equipment or adjacent racks. |
| Safety and Compliance |
Electrical safety, chemical handling, pressure safety, environmental controls, and applicable local regulations |
- Review pressure-vessel, plumbing, electrical, fire-protection, and occupational-safety requirements
- Maintain safety data sheets for all coolant additives and treatment chemicals
|
Complete a documented risk assessment covering leaks, spills, electrical exposure, hot surfaces, pressure release, and disposal of used coolant. |
| Scalability |
Future rack density, additional cooling loops, modular capacity, and expansion space |
- Plan for foreseeable increases in processor and accelerator thermal design power
- Reserve pipe capacity, electrical capacity, floor space, and control-system points
|
Use modular distribution and heat-rejection capacity where possible. Confirm that future expansion will not reduce redundancy or exceed hydraulic and thermal limits. |
| Acceptance Testing |
Factory testing, site testing, thermal-load testing, controls testing, and failure simulation |
- Test normal operation, peak load, pump failure, cooling-unit failure, power loss, leak alarm, and sensor failure
- Verify actual flow, pressure, temperature, and heat-removal performance
|
Approve the system only after measured results meet the design requirements and all operating procedures, alarm matrices, drawings, and maintenance documentation are complete. |