Why your UPS cannot support GPUs: the inertia problem no vendor tells you
If you migrated a traditional compute rack to an AI one, your UPS vendor probably told you the equipment supports the load. And in part it is true: it supports the nominal power. What it does not support are three combined physical effects that the data sheets do not document: the rotational inertia of the diesel generator, the rectifier response time under GPU 4 ms spikes, and the sizing of VRLA banks for repeated deep discharges. This article breaks down each one and proposes three criteria that DO work for validating a UPS before an AI deployment.
Why GPU racks break a commercial UPS
A server with CPUs consumes stable power: ripple between 200 and 500 watts absorbed without significant transients. A rack with 8 H100 GPUs consumes between 14 and 22 kW with spikes of up to 4 ms that double the nominal power during distributed training. That load creates three problems that traditional UPS handles poorly.
- IGBT rectifier: typical commercial UPS IGBT modules have response time of 2 to 6 cycles (33 to 100 ms in Mexico at 60 Hz). Between the GPU spike and the compensation, 33 to 100 ms have elapsed, during which the load falls to batteries.
- DC bus: commercial UPS from 100 to 500 kVA use a common DC bus between modules. Under non-linear step load, modules enter ringing between 50 and 80 Hz for 2 to 8 cycles, during which voltage regulation falls outside AT-001 tolerance (5%). The consequence is transient brownout that GPUs detect and disable compute cells for self-protection.
- VRLA bank: a UPS 200 kVA VRLA bank stores between 30 and 90 kWh at full autonomy (power factor 0.8, 80% depth of discharge). Under CPU loads, 8 to 12 minutes of autonomy is standard. Under GPU loads, autonomy drops between 30 and 50% for two reasons: deep spikes double the instantaneous current in 100 ms windows, applying double stress on each VRLA cell; training cycles distribute deep discharges in 60 to 90 second series with 30 to 45 minute recovery, not the flat pattern of the datasheet. This regime degrades each cell between 1.5 and 2 times faster than declared.
Why VRLA batteries fail under GPU loads
In production AI data centers (anonymous operators, ref. EIA 2024), a VRLA bank designed for 10 minutes of CPU autonomy delivers between 3 and 6 minutes under H100 at 80% depth. For 1.5 MW per rack, it is the difference between clean transfer to diesel and premature failure.
Why diesel generator rotational inertia matters
When a mains failure occurs, the UPS switches to batteries in milliseconds and transfers to diesel in 5 to 15 seconds. But generators have finite rotational inertia: the governor takes between 200 and 800 ms to adjust fuel injection. During that interval the frequency drops between 58.5 and 59.5 Hz (60 Hz system). H100 with strict regulation can trigger UVP (Under-Voltage Protection) and shut down preventively.
Three criteria to validate your UPS in one hour
For each you need less than one hour with your vendor or operations team; thresholds are clear and the criteria help prioritize replacements before commissioning AI racks.
- Request the datasheet of the power module (IGBT or SiC) from your vendor and measure response time under step load. A UPS suitable for AI racks must respond from 0% to 100% in less than 4 ms. If it responds in more than 4 ms, the drop is enough for the GPU to trigger internal protection.
- Measure effective autonomy under real GPU load. If it drops below 60% of the declared datasheet, the VRLA batteries are aging fast and must be replaced before deployment. Alternative: LiFePO4 (LFP) bank doubles cycling at 1.8 to 2.5 times the cost of VRLA but recovers 2 to 3 years of operational life.
- Request the technical sheet with Time to Frequency Recovery (TFR) from your diesel generator vendor. For AI racks, TFR must be less than 300 ms. Tier 4 Final generators of 2 MVA and above typically comply; those of 1 MVA or less rarely do. Consider a kinetic flywheel if TFR is between 300 and 600 ms; the CAPEX is recovered in 18 months by avoiding premature failures in AI racks.
When it IS worth replacing the UPS
Three combined conditions to amortize UPS replacement in 2 to 3 years:
- You have more than 50 kVA of GPU load planned in the next 24 months.
- Your current UPS is Tier 3 with response greater than 8 ms.
- Your diesel generators are 1 MVA or less.
When it is NOT worth replacing the UPS yet
Three cases where it is worth waiting before replacing the UPS:
- Planned loads below 50 kVA (racks with medium model inference, short fine-tuning or smaller batch loads).
- You have a Tier 4 Final diesel generator of 2 MVA or larger, which likely complies with TFR.
- You prefer using racks with AMD MI300X, which have more permissive internal voltage regulation than H100.
In those three cases, your existing UPS with IGBT modules can serve you one or two more years while you settle the scale-up decision.
Sources
- Uptime Institute — Tier classification, electrical and physical requirements for data centers — https://uptimeinstitute.com/
- ANSI/TIA — applicable standards (TIA-942-B-2023) — https://www.tiaonline.org/
- Vertiv — Three-phase UPS solutions and technical datasheets — https://www.vertiv.com/
- Schneider Electric — three-phase UPS catalog and backup architectures for data centers — https://www.se.com/us/en/product-category/three-phase-uninterruptible-power-supply.html
- NVIDIA H100 datasheet — PCIe power profiles, NVL and SXM5 modes — https://www.nvidia.com/en-us/data-center/h100/
Want to master this?
Noxtel Academy →