Investigation of electrical failures in data centers and other mission-critical facilities, led by an engineer who led electrical quality and field reliability engineering for a global hyperscale data center fleet from 2020 to 2025.
In a well-designed data center, the first failure usually costs no load at all: one side trips, the other carries the site, and the IT load never notices. That makes the investigation more urgent, not less. Redundancy has been consumed, the same cause may be present on the other side of the system, and there is intense pressure to restore the failed equipment, which is the evidence, as quickly as possible.
Establish control of the situation.
Then start the investigation.
Download the printable checklist and evidence log (PDF, 2 pages) →
The choice is rarely between getting the site back and preserving the evidence. A temporary re-feed from the healthy lineup can restore redundancy around the damaged section, provided the feeder is tested first, the relay settings on the breaker picking up the re-feed are revised for the new configuration, and the generator-transfer signalling is thought through for the temporary arrangement. Those decisions are engineering decisions, and they belong in the investigation record.
At about 6 a.m., one side of a facility substation transferred to on-site generation without losing load. The high-voltage breaker and the medium-voltage breakers on that side had tripped and locked out, and video showed smoke near the substation transformer.
The transformer differential relay had tripped on its C-phase restrained differential element. That relay accepts several three-phase CT sets plus neutral CTs, so the first task was mapping each set to its physical location from the protection drawings and the relay settings: the high-voltage side, two sets at the transformer’s secondary windings, and sets at each medium-voltage main breaker.
The event data then showed which CT sets carried fault current. The high-voltage set and one secondary-winding set did; the other secondary set and both switchgear main-breaker sets did not. That boundary places the fault between the transformer’s faulted secondary winding and its switchgear main breaker, in the bus transition from the transformer terminals to the building, before anyone opened an enclosure.
Inspection confirmed it, and found the cause. The installer had left out a gasket in the bus transition equipment, which let water in and degraded the equipment over time. Where the bus penetrated from the transformer enclosure it ran close to the grounded enclosure; the water bridged that gap and made it easier for electrical tracking to develop, which ultimately led to the fault.
Yes. The site ran on its redundancy, which means it has less of it until the cause is found and corrected, and a cause rooted in design, manufacture, or maintenance practice may be present on the other side of the system, or at other sites built to the same design.
Usually. Service can often be restored around the damaged equipment with a temporary re-feed while the failed section is held for inspection. The restoration plan should be reviewed with evidence preservation in mind, so no one has to choose between them.
Relay event reports and waveforms, because they overwrite on a rolling buffer. Then EPMS, BMS, UPS, and generator-controller logs, which age out on retention schedules, and any video of the space.
Direct principal access. Conflict check and retention letter typically within one to two business days.