J. Kissel for the Operator, Facilities, CDS/EE, Detector Engineering, Commissioning Teams Including but not limited to: D. Barker, R. Bork, J. Driggers, B. Gateley, C. Gray, J. Jones, K. Kawabe, P. King, J. Kissel, R. Kumar, N. Lecoeuche, G. Mansell, R. McCarthy, J. Oberling, H. Radkins, T. Sadecki, T. Shaffer, D. Sigg, J. Warner, B. Weaver, C. Vorvick Recovery Notes FRS Ticket 13311 Executive summary of the past 38 hours: - Total Time Between Observation Stretches: 37.85 hrs (spanning over 4 operators, and 6 operator shifts) 2019-07-24 03:56 UTC (Jul 23 2019 20:56:00 PDT): Report of lockloss from storm, computer system crash (LHO aLOG 50819) 2019-07-25 17:47 UTC (Jul 25 2019 10:47:00 PDT): Returned to Observing (LHO aLOG 50757) - Total Time of OBSERVATION TIME LOST from an actual FAULT: 12.53 hrs 2019-07-24 06:48:00 UTC first aLOG indicating that front-end computers have begun their normal recovery LHO aLOG 50762 2019-07-24 19:20 UTC Long-range dolphin Network Fixed LHO aLOG 50787 - The cause of the lockloss and subsequent computer systems crash at all stations has not yet been quantitatively proven. There are plenty of anecdotal coincident happenings (very bad lightning storm, heavy winds and rains, flickering of control room lights, flickering of "in town" lights, the mass storage room [where font-ends and timing system lives] going over to UPS for only 1 sec, a slight drop in power mains voltage), but yet nothing proven to be the cause. Research is on-going. - The dolphin network switch failure, covered in FRS Ticket 13311, is the only thing that failed due to, or was found to be broken after, this lightning-related event. This failure mode is suspected to be similar to that seen before at LLO (LHO aLOG 50788). The fix was to switch to another port on the dolphin switch (LHO aLOG 50786). Follow the ticket to see further course of action. - See notes below for the explanation of remaining time. Primary Causes of Time Loss During the Recovery: - The above mentioned hardware fault - three instances of humans not being aware/notified that disparate high-voltage power systems were not on (PSL PMC's PZT, ETMX's ESD Driver, and OMC's PZT), and - a single setting loss: the location of TMS pointing adjustment during PSL Power Up (LHO aLOG 50810) - All other time was time spent exercising normal recovery activity standard recovery from a computer outage, alignment recovery, red-herring chasing, and standard grief with the lock acquisition sequence, coupled with the sleep schedule of site staff. - For the three high-voltage power systems: the designed human notification system is the DIAG_MAIN guardian system. However, due to person power limitations (and lack of recent history of site-wide failures), the "checks" that this guardian does have not been updated to account for the improvements in the detector. This is now higher on our priority list, but it remains person-power limited. - For the one SDF setting that was problematic: the designed human notification system is the Settings Definition File (comparison) or SDF system. The setting *is monitored* in the SDF system, and likely threw a DIFF (humans have no way to trend what DIFF showed up when). The particular setting was one of many reported, over many front-end systems, to be different from the "safe" or "down" configuration. We chalk this human oversight to "unless one were one of the two people who installed the work-around to a commissioning problem (i.e. that we're hitting the edge of the transmon QPDs late in the lock-acquisition sequence, and you can't turn on the picomotors that were designed to fix the problem, so you steer the entire suspension in whatever bank you can find will work), then the recovery team -- a different, but equally "expert" collection of people -- will gloss over that problem as 'not a likely suspect.'" Our course of action [likely picked up by the operator and detector engineering teams] is to start understanding and accepting "safe" or "down" .snap differences and accepting them on a regular basis (monthly, perhaps weekly). The cadence will depend on the available person-power. Other Important Accumulation of Notes Many, Many Red-Herrings during the recovery: - 0.55 Hz oscillations in ASC system while holding at 2W, and playing around with gains therein. This was because the ASC system is not tuned for to stay a long time at 2W, and maybe because of cross-coupling between transmon QPD loops and input pointing (e.g. LHO aLOG 50798) - Suspicion that the ISS 2nd loop is somehow non-functional, or poorly aligned. It was fine. (LHO aLOG 50811) - ASC sensor dark offsets being wrong / out-of-date. They were fine. (not aLOGged, but QPD dark offsets were reset) - 9 MHz EOM driver noise being "higher than normal." It was fine. (not aLOGged, but was "bad" only intermittently during normal state transisiton) - Violin modes being too high to enter in to high power operations. They were loud, but not loud enough to be saturating PDs. (LHO aLOG 50821) - The EY dolphin network itself being broken/flawed. It was the switch. (see sequence of things ruled out in LHO aLOG 50786) - Alignment of the IMC optics and input periscope PZTs (LHO aLOG 50816), alignment of the OMs, being different from the past. They were different, but apparently inconsequentially different. - LSC REFLAIR PD's 27 MHz DEMOD LO was too low. Only just barely below threshold, not low enough to matter. (LHO aLOG 50809) - Evolution of POP18 and POP90 DRMI build-ups during IFO thermalization suspected to be too low. Yes, looks funking, but no different than normal. (LHO aLOG 50801) - Misalignment offsets for the ETMs. Only signs are different, and was errantly flipped a few weeks ago. (LHO aLOG 50824) Some of the things that would have been gotchas, but we remembered to do along the way in enough time that they didn't impede progress -- documented as "Recovery Notes": - PSL recovery (LHO aLOG 50779, LHO aLOG 50802, LHO aLOG 50776) - Restoring alignments of all optics to *slider* values, to a time *just before* lock acquisition LHO aLOG 50775 - Making sure that the test mass ring heaters are on with the right power settings *as soon as possible* LHO aLOG 50778, LHO aLOG 50777 - Check the functionality of all high voltage power supplies to avoid PSL PMC's PZT, ETMX's ESD Driver, and OMC's PZT problems - Run the seismic system's sensor correction configuration managers down from and back up to nominal configuration LHO aLOG 50795 - Verify that the photon calibrators are functional LHO aLOG 50817 - Verify that hardware injections are functional LHO aLOG 50818
DCPD saturation a few seconds right before lock loss and all of the signals on nuc2 dipping down.
J. Kissel and Rahul
This morning Jeff Kissel and I went through the SDF differences for ETMX and ETMY and found that the signs for the ETMX and ETMY top mass (M0) misalignment values were recently flipped. We trended the misalgiment values and found that these we last flipped on Feb 26, 2019 (after a DAQ restart LHO alog 47136). Given below are the channel name where the values of the misaligned state for the ETMX and ETMY are stored.
H1:SUS-ETMY_M0_TEST_Y_OFFSET Feb 26, 2019 (flipped from -16 to +8). 09 July 2019 flipped from +8 to -8
H1:SUS-ETMY_M0_TEST_P_OFFSET Feb 26, 2019 (flipped from -20 to +10). 09 July 2019, flipped from +10 to -10
H1:SUS-ETMX_M0_TEST_Y_OFFSET Feb 26, 2019 (flipped from +8.60 to -8.60). 25 July 2019, flipped from -8.60 to +8.35
H1:SUS-ETMX_M0_TEST_P_OFFSET Feb 26, 2019 (flipped from -9.77 to +10.41). 25 July 2019, flipped from 10.41 to -10.41
We have reverted the values for ETMY and accepted the values for ETMX. I am attaching a screenshot of the trends for the ETMX and ETMY, along with the SDF difference.
I found a few different settings that worked depending on the power in the IFO. At 2W of inpout power, ETMY Mode1 +60deg with -0.2 gain and Mode 6 +0deg 1.2 gain (could have been higher probably). As soon as the power began to increase, these ran away.
At 37W, I only had Mode1 +60deg and as much gain as I was comfortable with (~+1.2). This damped both modes.
ITMY Mode 13 is still a bit high, but since we are observing and teh filters are montitored in SDF I will leave it for now.
To assist robert in his 48Hz hunt. I changed the dipswitch settings of the CER AC unit (2A). this will take it out of the low noise setting. This work occurred at 1614UTC
Tagging PEM and DetChar
Many thanks to everyone's hard work to get us back to Observing!
SDFs:
Nicely done, all!
J. Kissel, T. Shaffer Running through guardians before heading to observe, we found the INJ_TRANS hardware injection guardian set to INJECT_KILL. We want these running (only run INJECT_KILL in GRB or GW events), so we set to INJECT_SUCCESS to get them running again.
After recovery MC1, MC3, and the IMC PZTs have all moved. We don't want to attempt to move them back now, but good to keep in mind.
J. Kissel Another thing not monitored because it's (a) guardian controlled, and (b) will change amplitude and frequency over the course of several lock stretches -- which means it's righfully monitored by SDF -- is the roaming high frequency calibration line drien by PCAL X. I noticed (from the DELTAL_EXTERNAL FOM on the wall) that the PCALX line looked pretty loud w.r.t. recent memory. Looking further at the screen on which these are defined -- I found (only because I remember what it used to be) that the frequency (1234 Hz) and amplitude (30000 ct) were before we made the changes to the guardian (see, e.g. LHO aLOG 49064). To fix, I just requested the HIGH_FREQ_LINES guardian to go to its INIT state, which restarted the sequence. The line is now at 4001.3 Hz as expected.
J. Kissel Having lost lock in the same "within a few seconds" timescale from doing nothing but rotate the rotation stage and adjust the "POWER SCALE" -- i.e. take the LASER_PWR -- guardian from POWER_2W to POWER_3W, I'm suspecting PDs falling off of useable area. We not the that IMC global position has moved quite a bit (at the ~10 urad level) from last week -- and since then we've had two major crashes (dolphin crash on Monday, and ) -- Here's a trend of the ISS QPDs. There's a major change on 2019-07-14 ~11:53 UTC (Sunday Jul 14 2019 04:53:00 PDT), but nothing recently of note. Attached is a trend from June 19 2019 to now (prior to June 19th, we're stable at the Jun 19th value since the start of O3).
IM1 and IM2 shifted in pitch, alogged here: alog 50571, and I called in, and talked to the operator, to make them aware.
Maybe someone has already looked into this but one of the first things we do in the INCREASE_POWER state is add an offset to TMS_X in pitch, to recenter the beam on the transmon QPDs
I had a look at one of the locklosses, and it looks like maybe the offset is too much, and when we turn it on we are falling off the transmon QPDs (see bottom plot of the attachment). I can have a look at other locklosses and see if that is not a fluke but maybe something to try, if it hasn't been done already, is to reduce this offset by ~ half?
THAT WAS IT! More details to come...
Whew!! ![]()
![]()
![]()
(The long term solution to this problem is to pico on the transmon so we don't have these sneaky hidden offsets in guardian)
J. Kissel, T. Shaffer, D. Sigg We found some suspicious diffs in the SDF for ECAT PLC2 (LSC-MOD_RF45_AM_RFSET, LSC-MOD_RF9_AM_RFSET -- which we subsequently found was guardian controlled in ISC_LOCK in states higher that we've been able to achieve -- so now real problem) -- but this led us down the ECAT rabbit hole to find that the LSC-REFLAIR_B_RF27_DEMOD now has its LO signal occasionally dropping threshold. A trend of the LO signal for this demod reveals that the LO level had been at 23 (dBm) for times back to August 2018, and then dropped on 2019-06-18 16:21:37 UTC (Tuesday, Jun 18 2019 09:21:37 PDT) to 22.7 (dBm) and then again on 2019-07-16 18:27:50 UTC (Tuesday, Jul 16 2019 11:27:50 PDT) to 21.9 (dBm). The threshold is 22 (dBm). Daniel suggests that this might be related to the balun replacement. Not enouh to worry about. "We'll fix it by lowering the threshold."
J. Kissel, T. Shaffer Back at the fight, we were able to get to ENGAGE_DC_VIOLINS for my first lock stretch this morning. Given the messages about the oscillation being seen in MICH, I had had the MICH P ASC gain (alone) reduced from -0.25 to -0.2. From trends courtesy of Niko, the DRMI power builds are essentially no different than before. Violin modes are "not unreasonable." ISS 2nd loop has also been checked for problems, and been cleared. Regardless, a power increase even from 2W to 3W (via adjusting the LASER_PWR guardian alone) triggers the 0.55 Hz QUAD Pitch resonance to ring up the ASC almost instantly. See attached, in which the power-up is at ~-120 sec.
Looking at the trend data, the NPRO switched off on July 23rd, 2019 at 10:33:46 PM local time (1247981644 GPS time) and was off for ~26 minutes.
Attached are plots of the NPRO, front end laser, and 70W amplifier outputs. The pre-modecleaner output and high voltage power supply output
for the same time period is attached (PMC.png and PMCKepco.png). The pre-modecleaner was not operational again for approximately another 4 hours
and 7 minutes. The FSS hadn't really stabilise until another ~37 minutes had passed (FSS.png). The ISS was not on until ~11.5 minutes after
the pre-modecleaner was restored (ISSLag.png).
The PSL was not functional until July 24th, 2019 at 2:57:42 AM local time (1247997480 GPS time), some 4 hours and 24 minutes after going
down.
Otherwise the PSL seems okay. It might be worth checking the performance of the servos at the next opportune maintenance Tuesday.
TITLE: 07/25 Owl Shift 07:00 – 15:00 (00:00-08:00), all times posted in UTC
STATE of H1: Corrective Maintenance
INCOMING OPERATOR: TJ
SHIFT SUMMARY: Still working to recover from power glitch yesterday. Locklosses continue at INCREASE_POWER, despite trying a few recommended changes by commissioners. Since my mid-shift post, I tried letting the violin modes ring down before increasing power (only got down to ~4.5 and lasted the longest of any attempts), realigned the OM’s to previous lock values, and tested that the IMC stays locked when powering up to 10 W (and that PWR_SCALE matches power).
LOG:
07:00 (00:00) Start of shift
07:16 (00:16) Starting initial alignment
07:36 (00:36) Initial alignment complete, starting re-lock
07:55 (00:55) Lockloss at CARM_TO_TR, looks like we lost DRMI while ramping down LSC-REFL_SERVO_IN2GAIN.
08:43 (01:43) Lockloss from INCREASE_POWER, after waiting at ENGAGE_DC_VIOLINS for ASC channels to converge
09:38 (02:38) Lockloss from INCREASE_POWER after reducing MICH_P gain by 70%
13:04 (06:04) Lockloss from INCREASE_POWER after waiting for violins to go down a bit
14:47 (07:47) Lockloss from INCREASE_POWER after aligning OM’s to previous config (to reduce AS_A_NSUM)
15:00 (08:07) End of shift
Once we regained Nominal Low Noise, one of the OBSERVE SDF DIFFs was in the OM3 alignment offsets. Niko comments about that he restored to previous lock's values, but those values appear to be quite different from what it's been for months. While in NLN, I set the offset's tramp to 360 seconds (6 minutes), and restored the offsets VERY SLOWLY. Then reverted the ramp time to its normal 2 seconds.
Corey called shortly after 14:00 PDT and informed me that the TCSy laser was off. Looking at the MEDM screen, the RTD/IR SENS. ALARM was red and the laser was off. I went to the LVEA and reset the laser at the front panel; I also took a picture of the front panel before I reset it, I'll attach it as a comment (it's on my phone at the moment). This cleared the temp alarm and the laser restarted without issue.
Trending back with ndscope (see 1st attachment), the laser tripped off around 11:41am PDT. The laser temperature was holding steady at the time of the trip, so the laser did not overheat. While the laser was off, the chiller setpoint was stepping from 20 °C to 21 °C in 0.1 °C steps; it went through this a couple times, as seen in the attachment. Seeing this, I went out and checked on the chiller and found it was reporting a water temperature of 22.3 °C; the setpoint for the laser after restarting was 20.9 °C. The 2nd attachment shows the water temp reported by the chiller starting . At the time of the trip the chiller was reporting a water temp of 22.0 °C, and is currently reporting 22.1 °C (There appears to be a descrepancy between the chiller setpoint and the water temp reported by the chiller, which I have never noticed before. I.E. before the trip the setpoint was just over 20 °C, while the chiller was reporting a water temp of 22.0 °C. Is this normal? Or could this be part of the cause of the trip?). At this point it is unclear what caused the laser to trip. Investigation continues.
I should also note, while I was looking at the TCSy chiller, I ran into Jeff B. He pointed out something interesting, the small display screen on the TCSx chiller has gone blank. The chiller is still running and the TCSx laser appears to be working fine, so it appears the display has simply quit working (I've never seen them blank before (like one would expect when a display goes to sleep), everytime I go to check on the chillers they displays are active).
Edited at 15:35 PDT as I hit the POST button a little too early.
A couple pictures. The first is the front panel for the TCSy CO2 laser. Notice that the IR Fault light is not lit, but the Temp alarm is. I seem to remember that when the viewport IR sensor trips, the IR Fault light flashes but does not hold, but I don't think the Temp alarm accompanies this. So I think this rules out the viewport IR sensor.
The 2nd attachment is the blank display on the TCSx chiller, while the 3rd is the active display on the TCSy chiller.
Entered FRS 13309 for this trip.
The chiller screen has been blank since January when Patrick reenabled the serial communication (alog46289).
The discrepencies in the tempuratures is interesting. I think this may be a good lead because these TCSY lock losses have become more frequent when we put the spare chiller in.
I think TJ might be on to something here. I trended the power and temp of the TCSy CO2 laser, and the temp reported by the TCSy chiller, through the month of January 2019, bookending the chiller swap that occured on Jan 15 2019. After the swap the laser temp increased from ~24.5°C to ~25 °C. It hovered there for about 3 days or so, then jumped up to ~ mid-26 °C, with jumps up to over 27 °C. In addition to this, the laser temperature became much more erratic after the swap, and that behavior has continued to today. It should be noted that the TCSx CO2 laser sits at a laser temperature of ~24.3°C, around where the TCSy CO2 laser temp used to sit before the chiller swap. Next step is to look back into 2018 to see if the stable pre-chiller-swap TCSy laser temperature behavior is consistent or a fluke. Another interesting thing to check is if the TCSy CO2 PZT is moving more than the TCSx CO2 PZT. I seem to recall seeing an alog about this, I'll hunt it down and link it here if it's relevant.
More data mining.
1st attachment is the laser power and temperature, as well as the chiller setpoint and the temp reported by the chiller itself, for the 4th quarter of calendar year 2018 (Oct-Dec 2018). As can be seen, even though there are spikes, the temp is consistently in the low-24s °C, not bouncing around as seen after the chiller swap (in the plot I posted in the 3rd comment above). Also of note, the discrepency between the temp reported by the chiller and the chiller setpoint is present even on the original TCSy chiller. That said, this is further evidence that the spare TCS chiller doesn't seem able to hold the laser at a consistent operating temperature.
Also, I found the alog regarding PZT movement between the 2 TCS CO2 lasers; TJ noted this on July 10th in this alog. Taking this a step further, I plotted out the PZT signal for each CO2 laser for the month of January 2019 (as I did with the laser temp yesterday), as well as from Dec 2018 to Feb 2019 (inclusive; the cursors indicate the TCSy chiller swap). The CO2y laser PZT (PZTy from here on, easier to type) does move more than the CO2x laser PZT (PZTx), but the difference in PZTy movement pre- and post-chiller swap is less pronounced but noticable. This also holds when looking at the longer 3 month trend (interesting note: at times, PZTx moves more than PZTy, but it's rare). I then looked at the PZT movement from Jan 1 2019 to now (final attachment, chiller swap again indicated by the cursors). PZTy in general is more active that PZTx and has been for some time. There are also stretches after the chiller swap where PZTy is moving more than its new norm, but this does not appear to be getting worse as time goes on.