Google Cloud named the cause of the us-central1 outage: technician disconnected all fiber optic cables in 13 minutes

Google Cloud published a preliminary report on the major outage in zones us-central1-b and us-central1-f on September 1, which lasted 4 hours 11 minutes. According to the document, the cause was a procedural error during planned router capacity expansion. A technician was replacing optical transceivers with denser ones, but the new modules proved incompatible with the data center equipment.
The work was supposed to be done one router at a time, with traffic rerouting and verification after each step. However, the maintenance manual had no instructions on the sequence of actions, and the required checks were missing. As a result, the technician disconnected all fiber optic connections on all devices within 13 minutes, isolating the affected facilities from the network.
At the peak of the incident, traffic drop reached 100%, virtual machines became unavailable externally and could not establish outgoing connections. Many Google Cloud services were affected: Compute Engine, Kubernetes Engine, Cloud Run, App Engine, BigQuery, Cloud SQL, Spanner, and others. The report notes that regional products suffered less due to automatic traffic failover.
The failure was detected almost immediately by monitoring systems, and engineers began diverting traffic away from the damaged infrastructure. Physical restoration of connections began after on-site specialists plugged the original transceivers back in. By 11:52 Pacific time, full service was restored.
Google apologized to customers and listed measures to prevent similar incidents. These include mandatory verification of light on each disconnected fiber, transitioning update processes to fully automated orchestration, and implementing an emergency stop system when active cables are accidentally disconnected. It also plans to accelerate automatic traffic redirection from faulty zones from 19 to 5 minutes.
According to the report, the incident affected only part of the us-central1-b and us-central1-f zones, and multi-zone deployments with redundancy in other zones of the region largely continued to operate.
Primary source: status.cloud.google.com ↗