david.barker@LIGO.ORG - posted 08:23, Thursday 30 January 2020 - last comment - 10:51, Thursday 30 January 2020(54813)
restarted external alert system
While H1 is out of lock I restarted the lv_alert external alert service on the ext-alert virtual machine. This is the second time I have had to do this this week.
Comments related to this report
thomas.shaffer@LIGO.ORG - 10:51, Thursday 30 January 2020 (54817)OpsInfo
This looks to be the known issue of the LVAlert subscriptions silently disappearing (I can't seem to find the link to where the bug is reported). This used to happen once in a blue moon, but lately it seems to be more frequent.
I talked with Jonathan and I'll write a script to monitor the heartbeat and then restart the systemd process if it flatlines. This should automatically take care of the resubscribing to the LVAlert nodes that we need.
In the mean time, Operators if you see the "GraceDb Query Failure" icon on the ops overview, please see the wiki in the follow link for instructions to restart the LVAlert service. https://cdswiki.ligo-wa.caltech.edu/wiki/RestartingLVAlert
This looks to be the known issue of the LVAlert subscriptions silently disappearing (I can't seem to find the link to where the bug is reported). This used to happen once in a blue moon, but lately it seems to be more frequent.
I talked with Jonathan and I'll write a script to monitor the heartbeat and then restart the systemd process if it flatlines. This should automatically take care of the resubscribing to the LVAlert nodes that we need.
In the mean time, Operators if you see the "GraceDb Query Failure" icon on the ops overview, please see the wiki in the follow link for instructions to restart the LVAlert service. https://cdswiki.ligo-wa.caltech.edu/wiki/RestartingLVAlert