[Tech Trend] The Breaker Is Fine, So Why Did the Power Go Out? – Causes and Countermeasures for Data Center Voltage Sag
Key point
Even without a blackout, Voltage Sag can halt data center facilities and GPU servers.
Details
In data centers, even with no history of blackouts, instantaneous Voltage Sag can cause chillers, cooling water pumps, and GPU server racks to shut down in a chain reaction. Even when the circuit breaker and incoming voltage appear normal, if the voltage momentarily drops due to lightning or grid fluctuations, IT equipment effectively perceives this as a supply interruption.
Voltage Sag is a phenomenon in which voltage temporarily drops to 10~90% of normal voltage, lasting 0.5 Cycle (about 8ms) to under 1 minute. Since the voltage does not reach 0V, it is not classified as an Interruption (blackout), but it can be sufficiently fatal to servers and communication equipment that operate on millisecond timescales. Even among the same category of power anomalies, Sag, Swell, Transient, Harmonics, and Interruption differ in waveform and countermeasures.
The causes fall broadly into two categories. On the external grid side, lightning strikes, ground faults, large nearby load switch-ins, and protective relay operations create Sags, while on the internal facility side, large motor startups, transformer inrush current, and ATS/CTTS switchover tests cause instantaneous voltage drops. In other words, Sag is not simply a failure, but a physical response produced as a large power system operates normally.
The reason this remains a problem even with a UPS in place is clear.
- Some Sags fall within the UPS's input tolerance range, so the UPS does not switch to battery mode.
- Even if UPS output is maintained, PSUs at the end of the rack or certain equipment can be vulnerable to voltage fluctuations outside the ITIC curve.
- In Eco mode, the momentary transition between commercial power and the inverter becomes an even greater risk.
In failure analysis, Sag should be suspected when the following pattern appears.
- There is no history of blackouts or generator operation.
- UPS alarms are absent or very minor.
- Only specific vendor equipment or specific lines selectively reboot.
- A voltage drop at the same timestamp is confirmed in PQM (Power Quality Meter) logs.
The response should be proactive control rather than after-the-fact recovery. Critical loads should be protected with a Double Conversion (Online) UPS, and a DVR (Dynamic Voltage Restorer) should be considered for ultra-sensitive loads. On the operational side, continuous waveform-level power quality monitoring, standardization of Sag thresholds, and cross-analysis of PQ logs, equipment logs, and UPS logs are needed. kt cloud is enhancing the reliability of failure root-cause analysis by precisely measuring and analyzing Sags that occur without any history of blackouts, using ultra-high-speed PQM monitoring.
Ultimately, Tier, redundancy, and Availability are merely result indicators. True reliability comes from the operational capability to observe and control voltage waveforms down to the 0.001-second level.