[Jenne, Cheryl, Rick]
I glanced at the spectrum on the wall, and noticed that the PcalY lines weren't on even though we were observing. This is highly unusual, since the calibration lines should always be on when we are in Observation.
I opened up the CAL_LINES overview screen (sitemap -> top cal button -> CAL_LINES) and saw that H1:CAL-PCALY_OPTICALFOLLOWERSERVOOSCILLATION was high at about 9.5 (text on the screen says that it should be below 0.01) and had a big red box around it. I called Rick for advice, and he suggested the ol' level-0 thing to try: turn it off and on again.
Cheryl and I took the IFO out of Observing, I turned off the PcalY optical follower servo (H1:CAL-PCALY_OPTICALFOLLOWERSERVOENABLE, the Loop Enable button on the PcalY overview screen), and then turned it back on. It seemed happy when it came back on, with the oscillation value back down to nearly zero, and also the H1:CAL-PCALY_OFS_PD_OUTMON wasn't railed flat. The Pcal lines in the DARM spectrum came back on, so we declared the IFO fixed, and went back to Observing.
By trending that oscillation channel, it looks like it was oscillating since about 21:48 UTC. This is about halfway through the time that the CDS team was at EY doing some corrective maintenance on the Beckhoff system down there (alog 48209 and comments). Daniel thinks that perhaps it could be that this was the time when some cables were unplugged or something?
In the attached screenshot, you can see that the blue trace was bad for a long-ish time, including some time when the red trace (which is the OBSERVE bit) was 1. The tick down in the red trace is my taking the IFO out of Observe, you can see that the oscillation went away, then we immediately went back to Observe.
The data during that time should be entirely valid and good for astrophysical searches, but note that the calibration lines were not on, so some diagnostics may report bad / funny values.
In order to prevent this in the future, I propose that we add a "test" to DIAG_CRIT (and maybe also DIAG_MAIN) that will make a notification of a problem with either Pcal servo. DIAG_MAIN is useful, since it is what we look at on the wall, but it is on purpose ignored by the Observe bit. So, it needs to be in DIAG_CRIT so that we *can't* go to Observe without the Pcal lines being on and healthy. As long as there is enough of an error message in DIAG_CRIT that an operator can find and fix the problem, then maybe it doesn't have to be replicated in DIAG_MAIN.
Rick suggests that if the servo goes into oscillation again (say, if it wasn't caused by the Beckhoff work at EY), we could perhaps try turning the gain (H1:CAL-PCALY_OPTICALFOLLOWERSERVOGAIN) down by 3dB until we get a chance to remeasure the servo loop gain next Tuesday. Recall from alog 48206 that we measured the loop earlier today and it had a UGF of about 84 kHz with 50 degrees of phase margin, so should be plenty stable, but perhaps things 'settled' after we left?