Reports until 13:41, Thursday 03 January 2019
H1 DAQ (CDS)
david.barker@LIGO.ORG - posted 13:41, Thursday 03 January 2019 - last comment - 10:47, Friday 04 January 2019(46224)
h1nds1 froze, needed a system reboot

At 13:11 PST h1nds1 died. The console output was showing a kernel problem in the network stack. The only recourse was to reboot the computer using the front panel reset button.

This is a 2.6.35 kernel machine, but had only been running for 192 days so this does not look like a 208.5 day bug crash.

h1nds1 is the default nds server for the control room, so some nds clients may need to reconnect to the NDS server.

Comments related to this report
david.barker@LIGO.ORG - 13:44, Thursday 03 January 2019 (46225)

Looking at the other 2.6.35 kernel DAQ machines' uptimes shows:

h1dc0 192days (16 days to go)
h1tw1 211 days (+3 days over!)
h1broadcast0 199days (9 days to go)

I'm opening a WP to cover rebooting this machines next tuesday. h1tw1 may crash before then.

david.barker@LIGO.ORG - 14:14, Thursday 03 January 2019 (46226)

At 14:08 PST the daqd process on h1nds1 died with restransmission type errors. Logs do not show excessive data requests at the time. Monit was not monitoring the process for unknown reasons, so it did not get restarted.

Jonathan and myself got monit to start the process and it appears to be ok now.

david.barker@LIGO.ORG - 10:47, Friday 04 January 2019 (46234)

And lastly, trended data was not available after the DAQD restart, restarting NDS resolved this.