Reports until 18:50, Monday 17 June 2019
H1 CDS
david.barker@LIGO.ORG - posted 18:50, Monday 17 June 2019 - last comment - 07:53, Tuesday 18 June 2019(50009)
h1iscey kernel panic

The general cores on h1iscey have locked up. The models on this front end continue to run (pemey, iscey, alsey, caley) but their EPICS and DAQ data is invalid.

All front ends which share h1iscey's port on h1dc0 have invalid data, which should clear once mx_stream is going again on h1iscey.

I am starting the reboot process.

Images attached to this report
Comments related to this report
david.barker@LIGO.ORG - 19:13, Monday 17 June 2019 (50010)

I have recovered h1iscey by remotely resetting it via its IPMI port. Procedure was:

  • take h1iscey out of the dolphin fabric
  • issue cpu reset via the IPMI management system

Once the mx_stream was restarted on h1iscey, the DAQ errrors on the other system cleared. 

Handing over to Jeff B for IFO recovery.

Details, taking h1iscey out of Dolphin fabric and verification:

david.barker@zotws6: ./dolphin_switch_port_status.sh h1iscey

h1iscey Switch_IP 10.101.0.94 Switch_port 3 is ENABLED

david.barker@zotws6: dolphin_disable_switch_port h1iscey

Disabling h1iscey Switch_IP 10.101.0.94 Switch_port 3

david.barker@zotws6: ./dolphin_switch_port_status.sh h1iscey

h1iscey Switch_IP 10.101.0.94 Switch_port 3 is DISABLED

Details, remote IPMI chassis reset

david.barker@zotws6: ipmitool -I lan -H 10.99.101.193  -U (user) -P (passwd) chassis status

Images attached to this comment
david.barker@LIGO.ORG - 07:53, Tuesday 18 June 2019 (50022)

FRS13078 opened to cover this issue.