TITLE: 11/29 Eve Shift: 00:00-08:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Lock Acquisition
OUTGOING OPERATOR: Jim
CURRENT ENVIRONMENT:
SEI_CONF state: WINDY
Wind: 13mph Gusts, 9mph 5min avg
Primary useism: 0.06 μm/s
Secondary useism: 0.24 μm/s
QUICK SUMMARY: IFO unlocked for no apparent reason ~2hours ago. Will try inital alignemt again, but green arm alignment seems way off.
TITLE: 11/29 Day Shift: 16:00-00:00 UTC (08:00-16:00 PST), all times posted in UTC
STATE of H1: Lock Acquisition
INCOMING OPERATOR: Camilla
SHIFT SUMMARY:
LOG:
22:00 Lockloss, no clear cause
Trying to relock, but the Y arm green refused to lock, ALS_YARM said "green input alignment off". Whatever the issue was, it was bad enough I couldn't even do initial alignment. Called Keita, and he got into the Y ALS screens, had to adjust H1:ALS-Y_PZT2_TRIG_THRESH_ON down to get some loop to come on. Handing off to Camilla now, trying to get ALS locked.
You might have noticed the strange gap in the wind trend ndscope covering the time h1seib3 was down (plot 3). This is because when h1seib3's general cores went offline, the mx_stream process (which sends the front end data to the DAQ) stopped running. This in turn caused problems for all the front ends which share the same ethernet port on the DAQ data concentrator. This is why most of the corner station frontend models went into DAQ data error (plot 2), and why it was critical to fix the problem quickly. One of the models so affected was the external EPICS data collector h1edc running on h1susauxh34 (plot 1). So all slow data was also frozen at the last received value.
TJ mentioned in his alog that the LOCKLOSS_SHUTTER_CHECK.py guardian code had problems reading the HAM6 GS13 channel during the time h1seib3 was down. This was another symptom of the DAQ data being corrupted on most of the corner station models, in this case h1isiham6's data.
This guardian node reads the fast GS13 channel H1:ISI-HAM6_BLND_GS13Z_IN1_DQ from the NDS, which would have been a repeat of the last 62.5mS of data (16Hz). Looking at this channel's raw data would appear to show noise. The attached minute trend plot shows this data repetition over the period h1seib3 was down.
Just a note from the future: the new CDS DAQ scheme utilizing a non-Myricom driver for the data concentrator (in full-scale testing from before start of O3), does not suffer this particular pathology. This has even been proven on L1 CDS, where the new scheme is running parasitically to an alternate data concentrator in parallel with the old one.
TITLE: 11/28 Owl Shift: 08:00-16:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Observing at 116Mpc
INCOMING OPERATOR: Jim
SHIFT SUMMARY: Quiet shift, locked for almost 8 hours. no issues to report.
Locked for almost 4 hours, wind is slowly calming down.
After a hard fought battle with wind and computer crashes, their efforts paid off and we are Observing! Range is slightly below nominal at ~114Mpc.
There were some SDF diffs for ISIITMX that I accepted since the model was recently brought back up, but perhaps the SEI team can look into why these are different.
LOCKLOSS_SHUTTER_CHECK node was in the SHUTTER_FAIL when I got here, and we were already at high power by the time I noticed. It looks like the check that it makes on the HAM6 GS13 didn't register a kick. I trended this and found that the GS13 channel was just noise during the time that the h1seib3 computer was down. ISC_LOCK was in NLN the entire time as well. I wrote off the shutter test failure to be related to the computer crash, though I'm a bit confused why this channel would be affected by the h1seib3 computer.
I then confirmed that our other fast shutter checking node, FAST_SHUTTER, had a successful test before going to high power. This is a check that it does every lock attempt, so it was reassuring to see that it passed the test for this lock try. With the fast shutter confirmed working, I took the LOCKLOSS_SHUTTER_CHECK node manually to HIGH_ARM_POWER, and then went to Observing.
Some of the differences that were accepted here turned off the CPS differential control for ITMX. Not sure what to do about this, but I think that all of the ISI controls guardians are written in such a way that they won't properly recover the CPS differential controls, blends or sensor correction if an ISI is restarted. The only way to fix this is to init all of those guardians for the chamber. It took until today to find this, because this log was not marked in anyway that got my attention. Unclear if there is anyway to tell if this affected our stability over the weekend, but likely it would have affected our robustness during earthquakes, probably wind, too. Also unsure of the interaction with the microseism focused seismic configurations, but none of those transitions were broadcast either.
TITLE: 11/28 Eve Shift: 00:00-08:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Corrective Maintenance
INCOMING OPERATOR: TJ
SHIFT SUMMARY: Still unlocked from SEIB3 computer crash. Now looking better. Wind high.
LOG:
TITLE: 11/28 Owl Shift: 08:00-16:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Corrective Maintenance
OUTGOING OPERATOR: Camilla
CURRENT ENVIRONMENT:
SEI_CONF state: WINDY
Wind: 28mph Gusts, 21mph 5min avg
Primary useism: 0.11 μm/s
Secondary useism: 0.29 μm/s
QUICK SUMMARY: Higher winds and a bit of snow. Currently relocking at increase power after h1seib3 crash.
J. Kissel We're continuing to have trouble acquiring lock after the SEIB3 computer crash (LHO aLOG 53533). I'm reasonably confident it's because of the obscene amounts of constant wind, but I'm nervous that it may also not be helping that we've decreased the effective range of the H1 SUS ETMX UIM/L1 actuators with the recent "permanent" change to turn on one stage of low pass (LHO aLOG 53528). In order to make the call as to whether you *need* to revert it, watch a live time series of the ETMX L1 requested DAC output on an ndscope, over a ~80 second timescale during any acquisition sequence: ndscope -s H1:SUS-ETMX_L1_MASTER_OUT_UL_DQ H1:SUS-ETMX_L1_MASTER_OUT_LL_DQ H1:SUS-ETMX_L1_MASTER_OUT_UR_DQ H1:SUS-ETMX_L1_MASTER_OUT_LR_DQ If these channels are regularly spiking above 2^18 ~ 131000 counts at various points during the acquisition sequence, then we need to revert. Here's how one would revert it: (1) Upon the next lock loss, hold the ISC_LOCK guardian in DOWN. (2) Change the coil driver state, using the following terminal command: caput H1:SUS-ETMX_BIO_L1_STATEREQ 1.0 (3) Resume lock acquisition as normal, but (4) Make it "permanent" by editing line 50 of the ALS_DIFF guardian code: change ezca['SUS-ETMX_BIO_L1_STATEREQ'] = 2 # Switched to requesting "one low pass on" from "no low passes on" (JSK 2019-11-27) to ezca['SUS-ETMX_BIO_L1_STATEREQ'] = 1 #Jeff's change didn't work, back to the drawing board. (5) Save the changes, and load the modified ALS_DIFF guardian code. (6) Once you get to nominal low noise, you'll have SDF DIFFs to accept; they'll look like the opposite of the attached screen shot, referencing "FM2" and the above mentioned BIO state request channel. (7) aLOG that you've reverted the settings.
h1seib3 has suffered a general core crash, meaning the models are still running (h1iopseib3, h1isiitmx and h1hpiitmx) but the DAQ data from itself, and most of the other SUS, SEI and PSL front ends are invalid (see attached). The only recourse is to reboot h1seib3. In preparation for this, I took this machine out of the Dolphin fabric (which killed the lock). Jeff is in the MSR taking a photo of the console and rebooting h1seib3 via the front panel reset button.
I've opened FRS13888
Here?s a picture of the rack ?KBM? console before we restarted the computer, as instructed by Dave.
The computer was powered down by holding the power button, but instead of a clean off-then-on the leds lit up immediately, which suggests the power button may have gotten wedged by a misaligned bezel (and I was not able to ping the machine even after several minutes). Jeff went back into the MSR to free the button, and I quickly re-disabled the IX dolphin port in case the computer had in fact started to reboot, in which case the second reboot would have crashed the entire corner station.
At this point in time h1seib3 has booted and its models are running, but the IOP model has a negative IRIG-B excursion so all IPC and DAQ data are bad. It wont be until the timing is correct will I know if the Dolphin port is still disabled, in which case I will re-enable it.
h1iopseib3 timing came good, its IPC and DAQ data is now good.
To complete the cleanup, I pressed the DIAG_RESET buttons on all models and cleared h1dc0's CRC counters. The CDS Overview is now green again.
Handing over to Camilla and Jeff for IFO recovery following a SEI-ITMX model reboot.
Post-Dated Entry: Here is the CDS Overview at the time of the crash:
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler. We have this implemented at LLO - see OpsWiki-Front-ends. As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here). The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS. But we want a lot of full-scale testing done first, so after O3.
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler. We have this implemented at LLO - see OpsWiki-Front-ends. As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here). The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS. But we want a lot of full-scale testing done first, so after O3.
[J.Kissel, T.Mistry]
After the calibration measurement were complete, I went to the X end and powered on and off the NCAL optical encoder system in 20 intervals. The interferometer was in Nominal Low Noise (NLN) at the after the calibration measurements. The motivation of this was to investigate whether the encoder would cause glitches and having known on and off times within the same lock stretch may indicate if this is the case. The first figure attached shows the encoder signal from the channel H1:CAL-NCALX_ENCODER_VELOCITY_OUT_DQ. When the encoder system is powered off, the channel measures 80 +/-1 DAC counts. When the encoder system is turned on, the channels measures 6932 +/-1 (approx 4V). A follow up invesitgation, in addition with investigation from LHO alogs 53503,53441 and 53396. A time log of the activities as taken by J.Kissel are as follows:
1258926919 NLN (Timesh arriving at end station, Jenne feedforward off time stops, Robert still somewhere in/around the YVEA with Wifi and phones ON doing camera recordings of test mass glint, wind fence crews are driving around at EY, audible within the YVEA).
21:55:01 UTC
1258927033 Timesh calls for first go in
21:56:55 UTC
1258927169 +/-15V power supply to mini-field rack Flipped ON. Begin first 20 minutes of ON data.
21:58:59 UTC
(XVEA Lights were left ON)
1258927372 lights OFF now
22:02:34 UTC
1258928389 Timesh begins to head in to turn off NCAL power supply (with lights remaining off)
22:19:31 UTC
1258928508 power supply flipped OFF
22:21:30 UTC
1258928589 Timesh calls back to say he's done turning OFF NCAL.
1258929539 Timesh heads back in to turn NCAL ON
1258929649 NCAL powered ON
22:40:31 UTC
1258930068
22:47:30 UTC
Keita OKs leaving the NCAL ON, as long as we are committed to showing that it doesn't matter.
Timesh goes in to garb room to clean up.
1258930168 Timesh leaves EX
1258930425 IFO handed over to Robert for his tests.
In case it wasn't clear from Timesh's entry: We have left the NCAL system *powered ON,* but not spinning. The data was supremely glitchy today, in general, because if the current wind storm. Thus, it was difficult to make any sound conclusions about whether we were able to reproduce the elevated glitch rate that was reported previously when the NCAL system was similarly powered on, but not spinning (see LHO aLOG 53396). So, we're leaving it powered on but not spinning for a few days to (a) hopefully get some data while the environment is more typical (i.e. quiet) (b) gather days-long stretches of data to be sure leaving it powered on doesn't adversely affect the CW search group. If exonerated, then we don't have to drive to the end station every time we want to use the thing. Stay tuned!
See LHO aLOG 53666 for a first look at this data using Fscans. Further study is needed, but the first look indicates some additional lines or increased line artifacts around this time.