TITLE: 08/17 Owl Shift: 07:00-15:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Observing at 108Mpc
OUTGOING OPERATOR: Ed
CURRENT ENVIRONMENT:
Wind: 21mph Gusts, 16mph 5min avg
Primary useism: 0.06 μm/s
Secondary useism: 0.09 μm/s
QUICK SUMMARY: range is more variable tha usual, still, locked in Observe
TITLE: 08/17 Eve Shift: 23:00-07:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Observing at 104Mpc
INCOMING OPERATOR: Cheryl
SHIFT SUMMARY:
H1 very glitchy. BNS range has trended (porpoised) downwards since it was locked 8 hours ago.
LOG:
00:07 Saw a "hump" in DARM between 100-200Hz
2:30 Exchanged the Wind Speed screen for te VE Overview screen (temporarily) as remote vieweing had become compromised.
I am unable to phone the site either via land line or Verizon Wireless. Regardless, we need to have the Vacuum Site Overview available for remote viewing. Someone has replaced our screen with one showing the wind speed.
Don't forget there are two ways of remotely viewing the vacuum overview, the MEDM snapshot system:
https://lhocds.ligo-wa.caltech.edu/screens/png/H0_VAC_SITE_OVERVIEW_CUSTOM-current.png
and the image capture of the control room TV display:
https://lhocds.ligo-wa.caltech.edu/cr_screens/png/video1-1.png
The former should always be working no matter what is happening to video1 in the control room.
TITLE: 08/16 Eve Shift: 23:00-07:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Observing at 116Mpc
OUTGOING OPERATOR: Travis
CURRENT ENVIRONMENT:
Wind: 18mph Gusts, 11mph 5min avg
Primary useism: 0.03 μm/s
Secondary useism: 0.10 μm/s
QUICK SUMMARY:
Locked and Observing!
TITLE: 08/16 Day Shift: 15:00-23:00 UTC (08:00-16:00 PST), all times posted in UTC
STATE of H1: Observing at 111Mpc
INCOMING OPERATOR: Ed
SHIFT SUMMARY: With the BRS issues, smoking TVs, and phone system non-functionality, it has been an interesting day. Thankfully we are back to Observing.
LOG:
15:44 Peter out
16:26 Niko, Ethan to PCal lab
18:19 Ethan to PCal lab
18:31 Ethan out
18:37 Kyle to MY
19:00 Kyle back
19:38 Fil to MY
19:58 Lockloss, PSL frontend crash
20:10 Gerardo to LVEA taking measurements
20:20 Richard to CER power cycling PSL IO chassis
20:57 Jim to EX checking out BRS
22:51 Observing
I'm putting a summary of the things we tried with BRSX today here. It might be fixed, but we need to monitor it overnight.
First thing, Patrick and I remotely restarted the BRSX beckhoff computer. This didn't resolve the glitching.
Next, Hugh, Fil and I went to the end station. We checked nothing was hitting the box, the inside looked undisturbed. Fil and Hugh power cycled the 24V power to the beckhoff (not sure what all that involved) and I restarted the beckhoff computer again. Still glitching.
Later, I found that there were fluctuations on the reference spots on the CCD, as Jeff K put in the DCC earlier.
Hugh and I went to EX again, this time we pulled the power on the light source, then plugged it back in. While we were there, we also opened the BRS enclosure again and pulled the CCD of the top of the light pipe. After looking down into the light pipe, we noticed two pretty good sized, dead, bugs sitting on the viewport at the bottom of the light pipe. I couldn't get my phone to focus on them through the beamsplitter. Since we couldn't reach them and didn't want to take the light pipe off, we put the CCD back on.
The BRS is more or less back in it's nominal state and I've been watching it for the last hour. It hasn't glitched since being restored after this last incursion, but it's still pretty rung up. Looking at trends from earlier, it seems possible that the glitches don't appear when the BRS is moving above some threshold. Could also be that power cycling the light source fixed it.
We should still monitor the BRS for at least overnight, before trying to run with it again.
C. Hagedorn, M. Ross
After watching the video Jeff posted (G1901492), scanning through some of the glitches, and the fact that dead bugs were found, we believe that a live bug on the lens or vacuum window is the most likely culprit for the glitches. This will require intervention before this BRS can be trusted again.
It is quite unlikely that it is a light-source problem. Both patterns are illuminated from the same source and we were able to find transients that affected only one of the two patterns. Additionally the video shows asymmetric transient intensity variations which is very unlikely to come from the light source.
We believe that there's still a live insect inside the autocollimator in addition to the dead insect. Dead insects won't give frequent transients in both patterns. The transients in the Ref and Main traces and the distortions in the video are all consistent with a bug on the condensing lens, beamsplitter, objective lens, or window. This causes smaller and wider intensity dips than the bug that previously caused problems (40396) which was directly on the CCD.
With some help from TJ, I've added a simple test to DIAG_MAIN to help diagnose this in the future. It just looks at the 300mhz - 1 hz blrms for the BRS and if that channel goes above 20, DIAG_MAIN will report the BRS as noisy, see the last two lines here:
@SYSDIAG.register_test
def BRS_CHECK():
"""Check the two end station Beam Roation Sensors to make sure that
they are not in FAULT (as read from their Guardian status node state.
Also checks the status of the CBIT, to see if C code is live
"""
for end in ['X', 'Y']:
if ezca['ISI-GND_BRS_ETM{}_CBIT'.format(end)] == 0:
yield 'BRS {} C code has stopped'.format(end)
elif ezca['GRD-BRS{}_STAT_STATE'.format(end)] == 'FAULT':
yield 'BRS {} is in FAULT'.format(end)
if ezca['ISI-GND_BRS_ETM{}_R{}_BLRMS_300M_1'.format(end, 'X' if end == 'Y' else 'Y')] > 20:
yield 'BRS{} is noisy'.format(end)
On the attached trend, neither BRS goes above this value, or even close, in the last 30 days except BRSXwhen it went "buggy".
For several days now h1sussr3 SDF was in ALL mode (denoted by purple block on overview screen, means that all channels are being scanned for sdf-diffs and not only those denoted as monitored). This means the chance of an SDF diff is much higher. I have switched it back to MASK.
At 13:03 PDT all h1psl0 models stopped running. I found that one of the 16bit DAC cards was not seen on the pci bus.
We power cycled both CPU and IO Chassis, the missing card reappeared. The IOP model did not run due to the IX Dolphin card in the CPU, but not configured on the fabric *.
Powered down CPU and removed IX611 card. IOP still wont run, its configured to use the card.
Removed pciRfm=1 directive from h1ioppsl0 and installed new IOP model. It started but we had a negative IRIG-B excursion which took about 15 minutes to clear.
Started user models, all is running and Peter verified correct DAC drives on all models.
* a description of how we got here. During the 18 Sept 2018 upgrade of Dolphin from DX to IX, h1psl0 was upgraded along with all the other IPC'ed systems. We then entered the series of Dolphin crashes, which took the PSL down each time, we quickly took the PSL out of the Dolphin fabric(%)by simply disconnecting its IX cable and taking h1psl0 out of the Dolphin manager's configuration. This kept h1psl0 stable and no reboot was necessary at that time. Fast forward to today, and h1psl0 had not been rebooted for 332 days, which is why we are seeing these startup problems today.
(%) In consultation with Keita, Daniel and Peter, the Dolphin IPC was used for ISS third loop PD readout from the end stations, and DBB Jitter transmission to LSC. Neither were needed for O3.
Silver lining, I was thinking of removing the Dolphin card during the Oct commissioning break.
During the Oct break I'll remove the PCIE IPC receiver parts from ISS so it is not always in IPC error. I'll also remove the IPC sender from DBB.
J. Kissel, J. Warner In the spirit "lessons learned" -- how could we have caught the BRS EX problem (FRS Ticket 13417) sooner -- Jim reminds us that there are pre-made BLRMS EPICs channels of each BRS in the standard ground motion frequency bands, e.g. H1:ISI-GND_BRS_ETMX_RY_BLRMS_30M H1:ISI-GND_BRS_ETMX_RY_BLRMS_30M_100M H1:ISI-GND_BRS_ETMX_RY_BLRMS_100M_300M H1:ISI-GND_BRS_ETMX_RY_BLRMS_300M_1 H1:ISI-GND_BRS_ETMX_RY_BLRMS_1_3 H1:ISI-GND_BRS_ETMX_RY_BLRMS_3_10 H1:ISI-GND_BRS_ETMX_RY_BLRMS_10_30 The first attachment is an "All of O3" trend of these channels, and one can "clearly!" (in the "in retrospect, it's obvious!" sense) that the BLRMS in all bands gets constantly elevated above regular ambient ground motion on Aug 15 2019. Second attachment is a time zoomin, and an amplitude zoom out, of the past week. The BRS rarely has extended (> 1 hour) periods in which all bands are elevated 5 nrad_rms, and during this glitchy badness, the RMS soared well above 20 in all bands. It's not going to catch everything, but for starters, we suggest a DIAG_MAIN alarm that looks at the collection BLRMS channels, and if they're collectively seen above 20 for 1 hour, then we certainly will know to look there immediately. Also, we plan on adding these BLRMS channels to both the BRS screens and to create a monitor screen similar to the SUS actuator saturations screen, the Violin Mode Monitor screens, and the ISI stage performance matrices, where there's a color-coded report of the BLRMS level shown to report good and/or bad levels. Hopefully this'll help us four years down the road when this happens again!
J. Kissel, H. Radkins, J. Warner, [M. Ross remotely] FRS Ticket 13417 As we continue to investigate the problems with BRSX glitching, we've removed it's tilt subtraction from the sensor correction path at EX (see above cited FRS Ticket and LHO aLOG 51309). Once we did, we were able to successfully get through the entire lock-acquisition sequence as normal -- hence the reported observation stretch in LHO aLOG 51324 and 51325. As recommended by Michael we took a look at the *raw* raw BRS signal -- the fringe patterns as measured by the CCD. To do this requires killing the BRS c-code and logging in to Beckhoff ***, so we've done so and captured the video -- posted under G1901492. A frame-by-frame instantaneous plot of the fringes are on the left, and the corresponding long-term time-series of the calculated rotation signal from it is on the right. For the left plot, the lefter of the collection of fringes is the reference light reflection, and the righter is the reflection from the rotation sensor's suspended beam. One can see at the start of the video, that 2 glitches have happened already in the past 500 seconds, one at ~150 seconds and another, smaller at ~250 seconds. Then, right at the start of the video (within 2 seconds of it), a new glitch starts, and one can see *all* the fringes -- the reference fringes a little more clearly than the that beam fringes drop in amplitude by (0.9 - 0.85)/0.9 = ~5%. The fringe heights restore (and the glitch stops) at around 5-7 seconds. Best to download the movie and open in a player where you have the ability to control frame, since it's a subtle affect, best seen on repeat or flipping back and forth between YES glitch and NO glitch frames. Where Michael's previous experience seeing this behavior was an actual insect "bug" walking across the CCD, this appears to be a more systemic problem, of either - sensitivity drops in the CCD, or - light level drops from the LED light source. The investigation continues... *** Reminder that we wouldn't *have* to kill the BRS c-code to look at diagnostic signals, if the recent attempt at a code upgrade hadn't failed for unknown reasons LHO aLOG 50178. Another reminder -- this is the oldest BRS, installed in Aug 2014 -- see, e.g. LHO aLOG 13296.
VIDEO1 television has released its smoke and has been removed from the control room.
OPERATORS: This is the second CR FOM monitor that has burnt up in the past month (the last one was July 16), so something to keep an eye/nose on.
Investigating and restarting.
Peter, Richard, Dave:
Looks like we have lost a 16bit DAC, only 3 are visible.
First thing we are trying, complete power cycle of computer and IO Chassis. Luckily h1psl0 has been take off the Dolphin fabric, so not need to worry about that.
My first indication of this was that all the non-laser PSL MEDM screens froze. as I pulled up the FSS screen to check the reference cavity transmission and admired the smoke billowing from the monitor on the wall.
[The LHO H1 Team]
After a rough day and a half, we're back Observing!
Throughout this, we have reset our initial alignment setpoints, see first attachment.
Also, we have caught a tricksy situation with guardian that if we do an initial alignment and then the DRMI never has to run its down state (as happened just now, since we are locked on our first attempt!), the PRM M2 coil drivers will be left in their higher noise state. Because we do not believe that we can switch it while locked (and we're not willing to risk a lockloss), we are leaving the PRM with the slightly higher noise coil driver state for this lock, and modifying guardian so that this cannot happen again.
However, this means that next lock, we'll need to re-accept the correct values in SDF (the setpoint values in the second attachment are the correct ones that we'll have to re-accept).
As Jenne said, we have kept the inital alignment references that we changed yesterday 51307 so that we are acquiring with the spots in the center of the optics. This is working well for lock acquisition, and this morning we allowed all 3 of the ADS loops to converge before we moved the spots to their final locations. We moved them and waited for convergence before increasing the ADS gains, so this will add a little time in the ENGAGE_SOFT_LOOPS state. We've not added this in the guardian, but haven't loaded the change or tested the new code.
Tagging OpsInfo just in case we have a beautifully long lock stretch until 3a.
Note: This lock stretch will have a non-nominal SEI Configuration. We are using WINDY_NO_BRSX as the BRSx is still glitching.
Since the PSL computer just went down and we're out of lock, we've loaded the new code, so we'll get to excercise it shortly.
plots of the PZT, Camera 18 (PMC Output), and Camera 29 (IMC_IN), from 15 July 2019 to 16 Aug 2019.
So Camera 28 changes with the output of the PMC. In throey this beam should not change, however, the calculated centroid does, which is likely that the centroid of the PMC trans beam changed, however it can't be ruled out that the beam itself also changed.
Camera 29 sees the IMC_IN beam, and it can change with changes of the beam into the EOM, alignment changes of the EOM (who's mount loosesn over time, and it's been over a year since we checked these screws), and changes of the beam after the EOM, which can come from PZT power outages, and manual alignment.
Plots attached in the order of these dates of the changes seen on the PZT, camera 28 and/or camera 29:
We were having trouble locking with these "new" restored IO pointings, so we reverted them back to the previous lock values.
Don't use the camera trend for comparison purposes when the exposure setting and/or power changed.
Attached shows the 2-month trend of CAM29 channels. The data are plotted only when IMC_PWR_OUT>30W to avoid confusion.
CAM29 X and Y "jumped" for example by [+5.6, +8] pixels on July 17 (~day 29 on the attached, things are plotted only when IMC_PWR_IN>30W to avoid confusion), that's not a physical change, it's an artefact caused by changing the exposure setting of the camera. 2nd attachment shows that the exposure change alone can make this kind of jump when nothing else is going on.
In the first attachment, bottom left is the exposure setting (EXP), bottom right is the sum of all pixels (SUM). At SUM=1E7 nothing really saturates. Video image to the right of the plots was taken at 37W today when SUM was ~9.5E6, the historgram of the image shows that there was no saturation. But 2E7 would already be a serious saturation. Solid white disk with 200 pixels radius gives us SUM=3.3E7.
Looking at how badly this camera was saturated before July 17 and how it's not saturating any more, anything older than July 17 is not really usable for now-then comparison.
In addition, when the beam isn't really dark on the picture at the edge of the camera (aka the beam doesn't fit), even without saturation nor beam motion, the beam position estimate could change when the power and/or exposure setting change. And the beam doesn't fit, look at the image especially at the top edge. So July 17-July 22 data where the EXP was 3264 might be offset compared with post July 24 data where EXP is 2500.
Many people (Sheila, Jenne, Jeff K, Keita, Travis, Ed, others...)
It seems like there are at least three problems right now. The alignment problem seems nearly solved, but the difficulties we are having with ALS mean that it is hard to make progress.
Alignment progress today:
ALS:
ESD transition lockloss:
Sheila and I were chatting, and it looks like the BRS X problem (alog 51309) could have been the cause of the lockloss she mentions that happens after we've transitioned from ETMX. This is a delicate-ish spot, and it looks like the BRS had one of it's glitches / bumps / discontinuities a few moments before the lockloss. In the attachment, top trace is Xarm circulating IR power, second row is the BRS-corrected seismic signal, and third row is the ETMY PUM stage that saturated. It looks like ETMY (through the ASC) followed the BRS hiccup, saturated, and caused the lockloss. Likely this will not be a problem today, since we have taken the BRS X out of the signal chain.
Attached is a screenshot of the initial alignment references, displayed on JeffK's new summary screen for easy comparison later.
I've noticed the exit gate partially open a few times over the last few weeks, but hadn't noted it, but as Cheryl points out, it does make for issue of us not being able to prevent visitors from entering the site (when I've seen it open, I would enter the site by driving through--with the hope that when it is triggered open as I enter that it closes to its normal state).