Reports until 19:03, Wednesday 27 November 2019
H1 CDS
david.barker@LIGO.ORG - posted 19:03, Wednesday 27 November 2019 - last comment - 17:41, Thursday 28 November 2019(53533)
h1seib3 general core crash, models still running but most of Corner Station DAQ data bad

h1seib3 has suffered a general core crash, meaning the models are still running (h1iopseib3, h1isiitmx and h1hpiitmx) but the DAQ data from itself, and most of the other SUS, SEI and PSL front ends are invalid (see attached). The only recourse is to reboot h1seib3. In preparation for this, I took this machine out of the Dolphin fabric (which killed the lock). Jeff is in the MSR taking a photo of the console and rebooting h1seib3 via the front panel reset button.

Comments related to this report
david.barker@LIGO.ORG - 19:04, Wednesday 27 November 2019 (53534)

I've opened FRS13888

jeffrey.kissel@LIGO.ORG - 19:06, Wednesday 27 November 2019 (53535)
Here?s a picture of the rack ?KBM? console before we restarted the computer, as instructed by Dave.
Images attached to this comment
david.barker@LIGO.ORG - 19:19, Wednesday 27 November 2019 (53536)

The computer was powered down by holding the power button, but instead of a clean off-then-on the leds lit up immediately, which suggests the power button may have gotten wedged by a misaligned bezel (and I was not able to ping the machine even after several minutes). Jeff went back into the MSR to free the button, and I quickly re-disabled the IX dolphin port in case the computer had in fact started to reboot, in which case the second reboot would have crashed the entire corner station.

At this point in time h1seib3 has booted and its models are running, but the IOP model has a negative IRIG-B excursion so all IPC and DAQ data are bad. It wont be until the timing is correct will I know if the Dolphin port is still disabled, in which case I will re-enable it.

david.barker@LIGO.ORG - 19:25, Wednesday 27 November 2019 (53537)

h1iopseib3 timing came good, its IPC and DAQ data is now good.

To complete the cleanup, I pressed the DIAG_RESET buttons on all models and cleared h1dc0's CRC counters. The CDS Overview is now green again.

Handing over to Camilla and Jeff for IFO recovery following a SEI-ITMX model reboot.

david.barker@LIGO.ORG - 19:26, Wednesday 27 November 2019 (53538)

Post-Dated Entry: Here is the CDS Overview at the time of the crash:

Images attached to this comment
keith.thorne@LIGO.ORG - 17:41, Thursday 28 November 2019 (53553)
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler.  We have this implemented at LLO - see OpsWiki-Front-ends.

As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here).  The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS.  But we want a lot of full-scale testing done first, so after O3.
keith.thorne@LIGO.ORG - 17:41, Thursday 28 November 2019 (53554)
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler.  We have this implemented at LLO - see OpsWiki-Front-ends.

As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here).  The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS.  But we want a lot of full-scale testing done first, so after O3.