J. Kissel for the Operator, Facilities, CDS/EE, Detector Engineering, Commissioning Teams Including but not limited to: D. Barker, R. Bork, J. Driggers, B. Gateley, C. Gray, J. Jones, K. Kawabe, P. King, J. Kissel, R. Kumar, N. Lecoeuche, G. Mansell, R. McCarthy, J. Oberling, H. Radkins, T. Sadecki, T. Shaffer, D. Sigg, J. Warner, B. Weaver, C. Vorvick Recovery Notes FRS Ticket 13311 Executive summary of the past 38 hours: - Total Time Between Observation Stretches: 37.85 hrs (spanning over 4 operators, and 6 operator shifts) 2019-07-24 03:56 UTC (Jul 23 2019 20:56:00 PDT): Report of lockloss from storm, computer system crash (LHO aLOG 50819) 2019-07-25 17:47 UTC (Jul 25 2019 10:47:00 PDT): Returned to Observing (LHO aLOG 50757) - Total Time of OBSERVATION TIME LOST from an actual FAULT: 12.53 hrs 2019-07-24 06:48:00 UTC first aLOG indicating that front-end computers have begun their normal recovery LHO aLOG 50762 2019-07-24 19:20 UTC Long-range dolphin Network Fixed LHO aLOG 50787 - The cause of the lockloss and subsequent computer systems crash at all stations has not yet been quantitatively proven. There are plenty of anecdotal coincident happenings (very bad lightning storm, heavy winds and rains, flickering of control room lights, flickering of "in town" lights, the mass storage room [where font-ends and timing system lives] going over to UPS for only 1 sec, a slight drop in power mains voltage), but yet nothing proven to be the cause. Research is on-going. - The dolphin network switch failure, covered in FRS Ticket 13311, is the only thing that failed due to, or was found to be broken after, this lightning-related event. This failure mode is suspected to be similar to that seen before at LLO (LHO aLOG 50788). The fix was to switch to another port on the dolphin switch (LHO aLOG 50786). Follow the ticket to see further course of action. - See notes below for the explanation of remaining time. Primary Causes of Time Loss During the Recovery: - The above mentioned hardware fault - three instances of humans not being aware/notified that disparate high-voltage power systems were not on (PSL PMC's PZT, ETMX's ESD Driver, and OMC's PZT), and - a single setting loss: the location of TMS pointing adjustment during PSL Power Up (LHO aLOG 50810) - All other time was time spent exercising normal recovery activity standard recovery from a computer outage, alignment recovery, red-herring chasing, and standard grief with the lock acquisition sequence, coupled with the sleep schedule of site staff. - For the three high-voltage power systems: the designed human notification system is the DIAG_MAIN guardian system. However, due to person power limitations (and lack of recent history of site-wide failures), the "checks" that this guardian does have not been updated to account for the improvements in the detector. This is now higher on our priority list, but it remains person-power limited. - For the one SDF setting that was problematic: the designed human notification system is the Settings Definition File (comparison) or SDF system. The setting *is monitored* in the SDF system, and likely threw a DIFF (humans have no way to trend what DIFF showed up when). The particular setting was one of many reported, over many front-end systems, to be different from the "safe" or "down" configuration. We chalk this human oversight to "unless one were one of the two people who installed the work-around to a commissioning problem (i.e. that we're hitting the edge of the transmon QPDs late in the lock-acquisition sequence, and you can't turn on the picomotors that were designed to fix the problem, so you steer the entire suspension in whatever bank you can find will work), then the recovery team -- a different, but equally "expert" collection of people -- will gloss over that problem as 'not a likely suspect.'" Our course of action [likely picked up by the operator and detector engineering teams] is to start understanding and accepting "safe" or "down" .snap differences and accepting them on a regular basis (monthly, perhaps weekly). The cadence will depend on the available person-power. Other Important Accumulation of Notes Many, Many Red-Herrings during the recovery: - 0.55 Hz oscillations in ASC system while holding at 2W, and playing around with gains therein. This was because the ASC system is not tuned for to stay a long time at 2W, and maybe because of cross-coupling between transmon QPD loops and input pointing (e.g. LHO aLOG 50798) - Suspicion that the ISS 2nd loop is somehow non-functional, or poorly aligned. It was fine. (LHO aLOG 50811) - ASC sensor dark offsets being wrong / out-of-date. They were fine. (not aLOGged, but QPD dark offsets were reset) - 9 MHz EOM driver noise being "higher than normal." It was fine. (not aLOGged, but was "bad" only intermittently during normal state transisiton) - Violin modes being too high to enter in to high power operations. They were loud, but not loud enough to be saturating PDs. (LHO aLOG 50821) - The EY dolphin network itself being broken/flawed. It was the switch. (see sequence of things ruled out in LHO aLOG 50786) - Alignment of the IMC optics and input periscope PZTs (LHO aLOG 50816), alignment of the OMs, being different from the past. They were different, but apparently inconsequentially different. - LSC REFLAIR PD's 27 MHz DEMOD LO was too low. Only just barely below threshold, not low enough to matter. (LHO aLOG 50809) - Evolution of POP18 and POP90 DRMI build-ups during IFO thermalization suspected to be too low. Yes, looks funking, but no different than normal. (LHO aLOG 50801) - Misalignment offsets for the ETMs. Only signs are different, and was errantly flipped a few weeks ago. (LHO aLOG 50824) Some of the things that would have been gotchas, but we remembered to do along the way in enough time that they didn't impede progress -- documented as "Recovery Notes": - PSL recovery (LHO aLOG 50779, LHO aLOG 50802, LHO aLOG 50776) - Restoring alignments of all optics to *slider* values, to a time *just before* lock acquisition LHO aLOG 50775 - Making sure that the test mass ring heaters are on with the right power settings *as soon as possible* LHO aLOG 50778, LHO aLOG 50777 - Check the functionality of all high voltage power supplies to avoid PSL PMC's PZT, ETMX's ESD Driver, and OMC's PZT problems - Run the seismic system's sensor correction configuration managers down from and back up to nominal configuration LHO aLOG 50795 - Verify that the photon calibrators are functional LHO aLOG 50817 - Verify that hardware injections are functional LHO aLOG 50818