Displaying report 1-1 of 1.
Reports until 13:48, Thursday 19 January 2023
H1 CDS
david.barker@LIGO.ORG - posted 13:48, Thursday 19 January 2023 - last comment - 14:17, Thursday 19 January 2023(66888)
All models on h1susb123 stopped due to hardware memory error

At 13:00:03 PST all the models on h1susb123 stopped running. Soon afterwards the SWWD tripped on h1seib[1,2,3]

We were able to login and run dmesg, which showed a transient memory error which the real-time models were not able to ride through (dmesg output below with local times).

h1susb123 was recovered by:

stopping all models, disabling its Dolphin switch port, rebooting the frontend.

After the models had started again I cleared the DAQ-CRCs, did a full DIAG_RESET, untripped the SWWDs. We handed the system over to Ryan and Rahul to continue with the recovery of H1

[Thu Jan 19 13:00:03 2023] LIGO code is done, calling regular shutdown code
[Thu Jan 19 13:00:03 2023] h1iopsusb123: ERROR - An ADC timeout error has been detected, waiting for an exit signal.
[Thu Jan 19 13:00:03 2023] h1susbs: ERROR - An ADC timeout error has been detected, waiting for an exit signal.
[Thu Jan 19 13:00:03 2023] h1susitmx: ERROR - An ADC timeout error has been detected, waiting for an exit signal.
[Thu Jan 19 13:00:03 2023] h1susitmy: ERROR - An ADC timeout error has been detected, waiting for an exit signal.
[Thu Jan 19 13:00:03 2023] h1susitmpi: ERROR - An ADC timeout error has been detected, waiting for an exit signal.
[Thu Jan 19 13:00:35 2023] mce: [Hardware Error]: Machine check events logged
[Thu Jan 19 13:00:35 2023] {1}Hardware error detected on CPU0
[Thu Jan 19 13:00:35 2023] {1}It has been corrected by h/w and requires no further action
[Thu Jan 19 13:00:35 2023] {1}event severity: corrected
[Thu Jan 19 13:00:35 2023] {1} Error 0, type: corrected
[Thu Jan 19 13:00:35 2023] {1} fru_text: Card01, ChnB, DIMM0
[Thu Jan 19 13:00:35 2023] {1}  section_type: memory error
[Thu Jan 19 13:00:35 2023] {1}  error_status: 0x0000000000000000
[Thu Jan 19 13:00:35 2023] {1}  physical_address: 0x00000006cce1fec0
[Thu Jan 19 13:00:35 2023] {1}  node: 0 card: 1 module: 0 rank: 0 bank: 2 device: 14 row: 49977 column: 1016 
[Thu Jan 19 13:00:35 2023] {1}  DIMM location: P0_Node0_Channel0_Dimm0 DIMMA1 
[Thu Jan 19 13:01:07 2023] mce: [Hardware Error]: Machine check events logged
[Thu Jan 19 13:01:07 2023] {2}Hardware error detected on CPU0
[Thu Jan 19 13:01:07 2023] {2}It has been corrected by h/w and requires no further action
[Thu Jan 19 13:01:07 2023] {2}event severity: corrected
[Thu Jan 19 13:01:07 2023] {2} Error 0, type: corrected
[Thu Jan 19 13:01:07 2023] {2} fru_text: Card01, ChnB, DIMM0
[Thu Jan 19 13:01:07 2023] {2}  section_type: memory error
[Thu Jan 19 13:01:07 2023] {2}  error_status: 0x0000000000000000
[Thu Jan 19 13:01:07 2023] {2}  physical_address: 0x00000006ce63fec0
[Thu Jan 19 13:01:07 2023] {2}  node: 0 card: 1 module: 0 rank: 0 bank: 2 device: 14 row: 50075 column: 1016 
[Thu Jan 19 13:01:07 2023] {2}  DIMM location: P0_Node0_Channel0_Dimm0 DIMMA1 
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 0
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]: It has been corrected by h/w and requires no further action
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]: event severity: corrected
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]:  Error 0, type: corrected
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]:  fru_text: Card01, ChnB, DIMM0
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]:   section_type: memory error
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]:   error_status: 0x0000000000000000
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]:   physical_address: 0x00000006ce63fec0
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]:   node: 0 card: 1 module: 0 rank: 0 bank: 2 device: 14 row: 50075 column: 1016 
[Thu Jan 19 13:01:32 2023] {1}[Hardware Error]:   DIMM location: P0_Node0_Channel0_Dimm0 DIMMA1 
[Thu Jan 19 13:01:40 2023] mce_notify_irq: 1 callbacks suppressed
[Thu Jan 19 13:01:40 2023] mce: [Hardware Error]: Machine check events logged
[Thu Jan 19 13:01:40 2023] {3}Hardware error detected on CPU0
[Thu Jan 19 13:01:40 2023] {3}It has been corrected by h/w and requires no further action
[Thu Jan 19 13:01:40 2023] {3}event severity: corrected
[Thu Jan 19 13:01:40 2023] {3} Error 0, type: corrected
[Thu Jan 19 13:01:40 2023] {3} fru_text: Card01, ChnB, DIMM0
[Thu Jan 19 13:01:40 2023] {3}  section_type: memory error
[Thu Jan 19 13:01:40 2023] {3}  error_status: 0x0000000000000000
[Thu Jan 19 13:01:40 2023] {3}  physical_address: 0x00000006cfe5fec0
[Thu Jan 19 13:01:40 2023] {3}  node: 0 card: 1 module: 0 rank: 0 bank: 2 device: 14 row: 50173 column: 1016 
[Thu Jan 19 13:01:40 2023] {3}  DIMM location: P0_Node0_Channel0_Dimm0 DIMMA1 
[Thu Jan 19 13:02:12 2023] mce: [Hardware Error]: Machine check events logged
[Thu Jan 19 13:02:12 2023] {4}Hardware error detected on CPU0
[Thu Jan 19 13:02:12 2023] {4}It has been corrected by h/w and requires no further action
[Thu Jan 19 13:02:12 2023] {4}event severity: corrected
[Thu Jan 19 13:02:12 2023] {4} Error 0, type: corrected
[Thu Jan 19 13:02:12 2023] {4} fru_text: Card01, ChnB, DIMM0
[Thu Jan 19 13:02:12 2023] {4}  section_type: memory error
[Thu Jan 19 13:02:12 2023] {4}  error_status: 0x0000000000000000
[Thu Jan 19 13:02:12 2023] {4}  physical_address: 0x00000006d167fec0
[Thu Jan 19 13:02:12 2023] {4}  node: 0 card: 1 module: 0 rank: 0 bank: 2 device: 14 row: 50271 column: 1016 
[Thu Jan 19 13:02:12 2023] {4}  DIMM location: P0_Node0_Channel0_Dimm0 DIMMA1 
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]: Hardware error from APEI Generic Hardware Error Source: 0
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]: It has been corrected by h/w and requires no further action
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]: event severity: corrected
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]:  Error 0, type: corrected
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]:  fru_text: Card01, ChnB, DIMM0
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]:   section_type: memory error
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]:   error_status: 0x0000000000000000
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]:   physical_address: 0x00000006d167fec0
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]:   node: 0 card: 1 module: 0 rank: 0 bank: 2 device: 14 row: 50271 column: 1016 
[Thu Jan 19 13:02:33 2023] {2}[Hardware Error]:   DIMM location: P0_Node0_Channel0_Dimm0 DIMMA1 
[Thu Jan 19 13:02:49 2023] mce_notify_irq: 1 callbacks suppressed
[Thu Jan 19 13:02:49 2023] mce: [Hardware Error]: Machine check events logged
[Thu Jan 19 13:02:49 2023] {5}Hardware error detected on CPU0
[Thu Jan 19 13:02:49 2023] {5}It has been corrected by h/w and requires no further action
[Thu Jan 19 13:02:49 2023] {5}event severity: corrected
[Thu Jan 19 13:02:49 2023] {5} Error 0, type: corrected
[Thu Jan 19 13:02:49 2023] {5} fru_text: Card01, ChnB, DIMM0
[Thu Jan 19 13:02:49 2023] {5}  section_type: memory error
[Thu Jan 19 13:02:49 2023] {5}  error_status: 0x0000000000000000
[Thu Jan 19 13:02:49 2023] {5}  physical_address: 0x00000006d321fec0
[Thu Jan 19 13:02:49 2023] {5}  node: 0 card: 1 module: 0 rank: 0 bank: 2 device: 14 row: 50385 column: 1016 
[Thu Jan 19 13:02:49 2023] {5}  DIMM location: P0_Node0_Channel0_Dimm0 DIMMA1 
 

Comments related to this report
rahul.kumar@LIGO.ORG - 13:57, Thursday 19 January 2023 (66889)

We have reset the WD on both seismic and suspensions for BS, ITMY and ITMX and have recovered them all. The BS ISI's T240 got kicked and so it took some time for it come down below the threshold and then get to the nominal state (i.e Fully Isolated).

The IFO is now ready to be re-locked,

david.barker@LIGO.ORG - 13:57, Thursday 19 January 2023 (66890)

CDS overview just before recovery of h1susb123

Images attached to this comment
david.barker@LIGO.ORG - 14:16, Thursday 19 January 2023 (66891)

Opened Ticket FRS26570

david.barker@LIGO.ORG - 14:17, Thursday 19 January 2023 (66892)

Opened WP10932 to cover reseating/replacing h1susb123 DIMMs

Displaying report 1-1 of 1.