PEFGInvestigated. Engineered. Defensible.
Start a Conflict Check
Dwg PEFG-SVC-11 Entity PEFG, PLLC Rev A
Data Centers ยท Mission-Critical Power

Data Center Power Failure Investigation

Investigation of electrical failures in data centers and other mission-critical facilities, led by an engineer who led electrical quality and field reliability engineering for a global hyperscale data center fleet from 2020 to 2025.

Jay Prigmore II, Ph.D., P.E., Principal Engineer, PEFG

By Jay Prigmore II, Ph.D., P.E. Last reviewed · Prepared and reviewed by a licensed Professional Engineer

Redundancy changes the investigation

In a well-designed data center, the first failure usually costs no load at all: one side trips, the other carries the site, and the IT load never notices. That makes the investigation more urgent, not less. Redundancy has been consumed, the same cause may be present on the other side of the system, and there is intense pressure to restore the failed equipment, which is the evidence, as quickly as possible.

The first hours

Establish control of the situation.

  • Account for everyone, and walk the area for hazards.
  • Verify lockout/tagout and de-energization; secure the scene.
  • Notify the fire department if needed.
  • Stand up a cross-functional team to run restoration and investigation in parallel.

Then start the investigation.

  • Retrieve telemetry: relay event reports and waveforms, EPMS and BMS logs, UPS and generator controller logs, and system states before and after the event.
  • Photograph the equipment as it is, and retrieve any video.
  • Document visible damage; limit inspection to non-destructive steps. Opening doors is usually fine; removing or disturbing components is not.
  • If the event could lead to a claim, involve counsel at the start; privilege and notice questions are easier to settle before the investigation begins than after.

Download the printable checklist and evidence log (PDF, 2 pages) →

Restoring redundancy without destroying evidence

The choice is rarely between getting the site back and preserving the evidence. A temporary re-feed from the healthy lineup can restore redundancy around the damaged section, provided the feeder is tested first, the relay settings on the breaker picking up the re-feed are revised for the new configuration, and the generator-transfer signalling is thought through for the temporary arrangement. Those decisions are engineering decisions, and they belong in the investigation record.

Systems investigated

  • Facility substations, substation transformers, bus transitions, and busway, including enclosure sealing and water ingress.
  • Medium-voltage switchgear, paralleling gear, and transfer schemes.
  • UPS systems and batteries, including thermal events.
  • Generators, generator paralleling controls, ATS and STS equipment.
  • Low-voltage distribution: switchboards, busway, PDUs, and RPPs.
  • Protection and controls, including ZSI, bus and transformer differential, and “go-to-generator” logic.

A published case study: facility substation busway failure

From a case study Jay Prigmore authored for a co-presented 2024 IEEE IAS Electrical Safety Workshop tutorial. A teaching example, not a client matter.

At about 6 a.m., one side of a facility substation transferred to on-site generation without losing load. The high-voltage breaker and the medium-voltage breakers on that side had tripped and locked out, and video showed smoke near the substation transformer.

The transformer differential relay had tripped on its C-phase restrained differential element. That relay accepts several three-phase CT sets plus neutral CTs, so the first task was mapping each set to its physical location from the protection drawings and the relay settings: the high-voltage side, two sets at the transformer’s secondary windings, and sets at each medium-voltage main breaker.

The event data then showed which CT sets carried fault current. The high-voltage set and one secondary-winding set did; the other secondary set and both switchgear main-breaker sets did not. That boundary places the fault between the transformer’s faulted secondary winding and its switchgear main breaker, in the bus transition from the transformer terminals to the building, before anyone opened an enclosure.

Inspection confirmed it, and found the cause. The installer had left out a gasket in the bus transition equipment, which let water in and degraded the equipment over time. Where the bus penetrated from the transformer enclosure it ran close to the grounded enclosure; the water bridged that gap and made it easier for electrical tracking to develop, which ultimately led to the fault.

Standards the analysis draws on

  • IEEE C37.91 (transformer protection) and IEEE 242 (Buff Book) for protection and coordination.
  • IEEE 3006 series: reliability of industrial and commercial power systems, successor to the IEEE 493 Gold Book.
  • NFPA 70B and NFPA 70E: maintenance and electrical safety.
  • ASTM E1188 and ASTM E860: evidence preservation and examination.

Frequently Asked Questions

FAQ
No load was lost. Is the event still worth investigating?

Yes. The site ran on its redundancy, which means it has less of it until the cause is found and corrected, and a cause rooted in design, manufacture, or maintenance practice may be present on the other side of the system, or at other sites built to the same design.

Can the investigation run while the site is being restored?

Usually. Service can often be restored around the damaged equipment with a temporary re-feed while the failed section is held for inspection. The restoration plan should be reviewed with evidence preservation in mind, so no one has to choose between them.

What data should be pulled first?

Relay event reports and waveforms, because they overwrite on a rolling buffer. Then EPMS, BMS, UPS, and generator-controller logs, which age out on retention schedules, and any video of the space.

Retain PEFG for a Data Center Matter

Direct principal access. Conflict check and retention letter typically within one to two business days.