Reports until 13:04, Sunday 27 January 2019
H1 CAL
david.barker@LIGO.ORG - posted 13:04, Sunday 27 January 2019 - last comment - 18:26, Sunday 27 January 2019(46657)
corner station Dolphin crash, caused by spontaneous restart of h1oaf1 at 11:48 PST

Sheila, Dave:

The h1oaf1 machine restarted itself at 11:48PST, which in turn caused a corner station Dolphin crash. I have capture the logs and am in the process of restarting all the models. Unfortunately since h1oaf1 rebooted itself, its logs at the point of the restart have been lost.

Images attached to this report
Comments related to this report
david.barker@LIGO.ORG - 13:13, Sunday 27 January 2019 (46658)

The attached overview shows that h1oaf1's models are the only ones on the corner station dolphin'ed machines which do not have major errors.

The good news is that this is not a spontaneous Dolphin crash which is hopefully fixed by the Dolphin EEPROM upgrade on the 8th, unfortunately this resets the no-crashes-since-eeprom-upgrade timer after 19 days of running.

david.barker@LIGO.ORG - 13:16, Sunday 27 January 2019 (46659)

Opened FRS12217

david.barker@LIGO.ORG - 13:20, Sunday 27 January 2019 (46660)

Reminder that h1oaf1 is unique in that it is the only H1 front end computer which has no attached IO Chassis. Its IOP is a virtual-IOP, getting its timing from h1ioplsc0 (time master) via the Dolphin network.

david.barker@LIGO.ORG - 13:28, Sunday 27 January 2019 (46661)

All models back up and running. Although it may not have been necessary, I restarted h1oaf1's models (following h1lsc0) to ensure everything was started cleanly.

Handing over to Sheila for IFO recovery.

sheila.dwyer@LIGO.ORG - 18:26, Sunday 27 January 2019 (46662)

Thomas Vo, Sheila

These are the steps we are taking to recover. 

  • ran dark offstes script, since we noticed that dark offsets changed in the reboot. 
  • We also have many alignment sliders that were lost, so we looked at frs 12090, where betsy points out where we can find burt snapshots of alignment sliders.  We couldn't remember/no longer have aliases for burt, so we did this semi-manually (see comment in FRS, it would still be a good idea to make this easier for people)
  • Thomas moved the IMC PZT until the IMC locked, but the WFS aren't triggering. Got very confused by unloaded changes in the ASCIMC model and the IMC_LOCK guardian
    • Perhaps the triggering for IMC ASC has been changed from using IM4 TRANS SUM to NSUM, but the model hasn't actually been restarted yet.
    • Although the safe.snap had trigger thresholds which seem to have taken into account the nonsensical calculation above, the guardian was ersetting them to number between 30-60.  It turned out that the problem was that the numbers used for these thresholds had been updated in the guardian code, but that the guardian hadn't been updated.  So we have been running without IMC WFS for a while.  Since these settings don't need to change, we removed them from the IMC_LOCK down and MOVE_TO_OFFLINE (no idea why they would ever have been set in move_TO_OFFLINE, hardcoded in a second superfluous place). 
  • Thomas ran initial alignment, because we were in the middle of alignment when the computers crashed.
    • Earlier in the day I had lowered the gain for the X arm IR lock from 0.5 to 0.4 because it was oscillating while we were aligning the input beam, This seemed to work well
    • Thomas had to align SR2 by hand using AS_C
  • We tried to start locking with sensor correction off, we forgot that it is necessary after models go down to go back to the SEI_CONFIG node and re-request the current state so that it will turn sensor correction on again.
  • The POP A offsets that we manually tuned last week for DRMI ASC were not saved in SDF so the DRMI ASC brought us to a bad build up, this was fixed by restoring the offsets and saving them in SDF. 
  • In DRMI ASC, the beam splitter is controlled by AS45, this loop seems to be oscillating which it wasn't before. The oscillation goes away when it switches to 36.
  • We had trouble with the TR_CARM transition.  We were able to make it through the transition by not engaging the anti-boost FM4.  We've added that back in at the state "RESONANCE".  This problem is probably due to some change elsewhere, perhaps when we loaded the IMC guardian there was something unintended loaded.  Looking at the svn we don't see any thing that looks bad. 

We were locked at 2W and investigating ASC problems unrelated to the computer crash about 6 hours after the crash.