It seems like the node that casued us to drop out of observing was due to an automatic SEI_CONF change EQ>WINDY>EQ, maybe due to EQ being requested before we were fully transitioned to WINDY. Details follow:
- 00:38 SEI_ENV changed SEI_CONF to EQ for incoming 5.7M from Alaska. (for P-waves)
- 00:48 SEI_ENV changed SEI_CONF back to WINDY (main EQ waves have not yet hit).
- Maybe we need to wait more than 10 minutes in EQ or wait until the SEIMON alert is not active anymore before transitioning?
- 00:49 SEI_ENV changed SEI_CONF to EQ and we are kicked into COMMISSIONING. No SDF diffs that I saw so back to OBSERVING.
2020-03-02_00:49:09.429949Z IFO [OBSERVE.run] USERMSG 0: waiting for node: DIAG_CRIT
2020-03-02_00:49:09.274144Z DIAG_CRIT [RUN_TESTS.run] USERMSG 0: SEI_CONFIG_ACTIVE: HPI_BS_SC not ACTIVE
Looking at the logs, it looks like the HPI_BS_SC node may have stalled, which could be due to a known bug in the sensor correction parts of the simulink models. The ISIs have been restarted to address the bug, but the HEPIs may not have been. I need to get some help from Dave to determine if the HEPIs are running the most recent code or not, but it's possible we have a way to prevent this happening down the road.
This was actually a case where the worker for the Guardian node itself was terminated and then restarted. I don't why the worker timed out, but recovery was quick.