Following on from this morning’s failure on pve-1, Fido’s engineers were working on pve-2 to ensure a similar failure could not happen, however as they were disabling the kernel care services on pve-2 the hardware watchdog again kicked in and forced a reboot at approximately 12:15. Service was restored by 12:35.
This resulted in a brief outage to services hosted on pve-2 at Fido, including parts of the Fido Glide / Zimbra cluster, and redmail.com (note we operate distributed clusters in order to avoid total outages, but there would be service degradation whilst the hypervisor reboots).
Engineers are continuing to monitor services to minimise any further disruptions.
Apologies to customers inconvenienced by this brief outage.
