TITLE: 11/29 Owl Shift: 08:00-16:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Corrective Maintenance
INCOMING OPERATOR: Jim
SHIFT SUMMARY:
Locking Recovery After Computer Re-Boot: Not going so well.
LOG:
EX Saturations
13:54 Superevent S191129u
Looks like a general core crash, models remain running. Unlike earlier h1seib3 crash, in this case no other systems are affected and only DAQ data from SEI-B1 is down.
I've taken h1seib1 out of the Dolphin fabric, which caused lock loss. Ed is now rebooting the machine via the front panel reset button.
Opened FRS13892
Computer rebooted with no problems. No IRIG-B timing issue. All ADC and DAC cards seen.
Over to Ed for IFO recovery
14:12 SEI B1 computer went down and reported ITM trips, but H1 remains locked.
14:22 Contact made with Dave Barker - waiting on a return call
14:36 Dave is removing it from Dolphin to reboot which will likely caue lockloss
14:40 Lockloss due to computer work
14:42 H1IOPSEIB1 FE computer RESET
14:46 reset SWWD
While locked and Observing for 7hours, H1 has been increasingly glitchy with only 2 EX saturations
The wind is up in excess of 20MPH. There are no other obvious causes available on FOMs. Glitching ranges from half spectrum range (lower) to broad range.
TITLE: 11/29 Owl Shift: 08:00-16:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Observing at 115Mpc
OUTGOING OPERATOR: Camilla
CURRENT ENVIRONMENT:
SEI_CONF state: WINDY
Wind: 10mph Gusts, 8mph 5min avg
Primary useism: 0.03 μm/s
Secondary useism: 0.28 μm/s
QUICK SUMMARY:
Good job! And... thank you!
You might have noticed the strange gap in the wind trend ndscope covering the time h1seib3 was down (plot 3). This is because when h1seib3's general cores went offline, the mx_stream process (which sends the front end data to the DAQ) stopped running. This in turn caused problems for all the front ends which share the same ethernet port on the DAQ data concentrator. This is why most of the corner station frontend models went into DAQ data error (plot 2), and why it was critical to fix the problem quickly. One of the models so affected was the external EPICS data collector h1edc running on h1susauxh34 (plot 1). So all slow data was also frozen at the last received value.
TJ mentioned in his alog that the LOCKLOSS_SHUTTER_CHECK.py guardian code had problems reading the HAM6 GS13 channel during the time h1seib3 was down. This was another symptom of the DAQ data being corrupted on most of the corner station models, in this case h1isiham6's data.
This guardian node reads the fast GS13 channel H1:ISI-HAM6_BLND_GS13Z_IN1_DQ from the NDS, which would have been a repeat of the last 62.5mS of data (16Hz). Looking at this channel's raw data would appear to show noise. The attached minute trend plot shows this data repetition over the period h1seib3 was down.
Just a note from the future: the new CDS DAQ scheme utilizing a non-Myricom driver for the data concentrator (in full-scale testing from before start of O3), does not suffer this particular pathology. This has even been proven on L1 CDS, where the new scheme is running parasitically to an alternate data concentrator in parallel with the old one.
h1seib3 has suffered a general core crash, meaning the models are still running (h1iopseib3, h1isiitmx and h1hpiitmx) but the DAQ data from itself, and most of the other SUS, SEI and PSL front ends are invalid (see attached). The only recourse is to reboot h1seib3. In preparation for this, I took this machine out of the Dolphin fabric (which killed the lock). Jeff is in the MSR taking a photo of the console and rebooting h1seib3 via the front panel reset button.
I've opened FRS13888
Here?s a picture of the rack ?KBM? console before we restarted the computer, as instructed by Dave.
The computer was powered down by holding the power button, but instead of a clean off-then-on the leds lit up immediately, which suggests the power button may have gotten wedged by a misaligned bezel (and I was not able to ping the machine even after several minutes). Jeff went back into the MSR to free the button, and I quickly re-disabled the IX dolphin port in case the computer had in fact started to reboot, in which case the second reboot would have crashed the entire corner station.
At this point in time h1seib3 has booted and its models are running, but the IOP model has a negative IRIG-B excursion so all IPC and DAQ data are bad. It wont be until the timing is correct will I know if the Dolphin port is still disabled, in which case I will re-enable it.
h1iopseib3 timing came good, its IPC and DAQ data is now good.
To complete the cleanup, I pressed the DIAG_RESET buttons on all models and cleared h1dc0's CRC counters. The CDS Overview is now green again.
Handing over to Camilla and Jeff for IFO recovery following a SEI-ITMX model reboot.
Post-Dated Entry: Here is the CDS Overview at the time of the crash:
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler. We have this implemented at LLO - see OpsWiki-Front-ends. As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here). The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS. But we want a lot of full-scale testing done first, so after O3.
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler. We have this implemented at LLO - see OpsWiki-Front-ends. As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here). The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS. But we want a lot of full-scale testing done first, so after O3.