30+ UPS without maintenance: how a pharmaceutical plant in the State of Mexico went from 9 stops per year to zero

In industrial plants with decades of operation, it is common to find a large UPS inventory that, according to the maintenance team, work fine, but have never gone through a documented preventive maintenance protocol. The pharmaceutical plant in the State of Mexico that motivates this article had more than 30 UPS distributed across production areas, local server rooms, and process control panels, none with an audited log.

This article describes how the situation looked before the intervention, what risks were operating in silence, and what protocol was implemented to move from “the UPS work” to a documented, audited operation with clear metrics.

The case: 30+ UPS without a single protocol

The plant ran three production shifts with an electrical maintenance team that responded to failures as they appeared, not by schedule. Each UPS had its own age, brand, and condition: some with original batteries, others with replacements made by different vendors, and at least four with expired battery indicators that no one had formally recorded.

The first step was to inventory the 30+ UPS with a standardized record template: brand, model, capacity, installation date, operating hours, last documented intervention, physical battery condition, and capacitor bank condition where applicable.

Why industrial plants accumulate UPS without a strategy

The accumulation of UPS without a unified protocol is a frequent pattern in plants with organic growth:

  • Growth by project: Each new production line or server room brought its own UPS. No one consolidated the inventory.
  • Vendor rotation: Maintenance contracted with different companies over the years, without a unified log.
  • Focus on failure, not health: The electrical team responded when something stopped working, not when something was about to fail.
  • No auditable metric: There was no way to answer, with data, which UPS were in good condition and which required immediate attention.

The five risks of operating UPS without preventive maintenance

Once the 30+ UPS were inventoried, the audit identified five risks that were already operating in the plant:

  • Expired batteries: Approximately 25% of the inventory had batteries over five years old without replacement. The probability of failure during an actual outage was high.
  • Aged capacitors: Several UPS showed wear in the capacitor bank, with effective reduction of backup capacity.
  • Improperly sized load: Some UPS operated near 80% of their nominal capacity, with no margin for transient peaks.
  • No autonomy testing: No UPS had been tested under real load in the past 24 months. The documented autonomy was theoretical.
  • No effective redundancy: Critical UPS operated as N, not as N+1 (one backup component for each main component). A failure in one directly compromised the area it protected.

The protocol implemented: three layers

The protocol built on top of the inventory has three layers, each with defined periodicity and responsible party:

Layer 1: Monthly inspection (Internal)

  • Activities: Reading LED and display indicators on each UPS, visual verification (temperature, abnormal noise, smell) and recording of hours/events of battery transfer.
  • Responsible: Plant electrical technician.

Layer 2: Quarterly maintenance (Specialized vendor)

  • Activities: Internal cleaning of filters and fans, torque verification on battery connections, autonomy testing under controlled partial load, and calibration of voltage and current sensors.
  • Responsible: Manufacturer-certified vendor.

Layer 3: Annual audit (Manufacturer or independent third party)

  • Activities: Full autonomy test under nominal load, thermographic analysis of connections/batteries, and residual useful life assessment of each bank with prioritized replacement report.
  • Responsible: UPS manufacturer or certified third party.

Result after 12 months of protocol

After a year of operating under the protocol, the plant’s indicators changed across three dimensions:

  • Documented reliability: We moved from “the UPS work” to auditable metrics: percentage of UPS with batteries within useful life, percentage with up-to-date autonomy tests, MTTF (Mean Time To Failure) by brand and model.
  • Zero operational stops in the past 12 months: Unplanned failures and stops in critical production lines dropped from an average of 9 incidents per year to zero operational stops in the past 12 months, because quarterly interventions detected wear before it translated into failure.
  • Traceability for regulatory audit: COFEPRIS and other pharmaceutical sector regulators require documented evidence of electrical continuity. The protocol log is the evidence presented in audits.

Where to start if your plant is in this situation

If the scenario of UPS that work but have no protocol sounds familiar, the three first steps with immediate return are:

  • Inventory and standardized record: Without inventory, there is no protocol. One week of work delivers the baseline.
  • Identification of the 5-10 most critical UPS: Those that protect regulatory or continuous-process loads. Intervention priority over the rest.
  • Quarterly maintenance contract with certified vendor: A single vendor for the entire inventory, with a unified log and quarterly reports.

These three steps do not require significant initial investment and reduce the unplanned failure rate within the first six months.

The protocol as continuous operational protection

Preventive maintenance discipline in UPS is not an administrative task: it is continuous operational protection. In pharmaceutical plants, electrical continuity directly affects product integrity, batch validity, and the ability to comply with COFEPRIS requirements and current sanitary regulations.

A UPS fleet without a protocol is a silent vulnerability: it only manifests when an actual outage occurs, and by then the failure is already in the history. Building the protocol and the log is a decision that pays for itself with the first failure avoided.


Sources

[1] Uptime Institute — Data Center Resources: https://uptimeinstitute.com/resources

[2] Uptime Institute — Blog: https://uptimeinstitute.com/blog

[3] Wikipedia — Uninterruptible power supply: https://en.wikipedia.org/wiki/Uninterruptible_power_supply

[4] Wikipedia — Preventive maintenance: https://en.wikipedia.org/wiki/Preventive_maintenance

[5] Wikipedia — VRLA battery: https://en.wikipedia.org/wiki/VRLA_battery

[6] Wikipedia — Battery room: https://en.wikipedia.org/wiki/Battery_room

[7] IEEE Std 1188 (batteries): https://standards.ieee.org/ieee/1188/7230/

[8] Vertiv — UPS Services: https://www.vertiv.com/en-us/services/

[9] Wikipedia — Data center: https://en.wikipedia.org/wiki/Data_center

Also in Facility Management

← Back to categories