After ~50hrs, we have lost lock
1821 UTC
J. Kissel, L. Sun, Calibration measurements beginning as planned (see 48057)
Jonathan, Tanner, Carlos, Dave:
The external alerts system is up and running on its new dedicated machine. This is a VM installed by Carlos yesterday, called ext-alert. Jonathan installed the certificate authentication system. Many thanks to Tanner for help getting this going.
I made the same small change to {userapps}/cal/common/scripts/ext_alert.py as was done at LLO. This directory is locally checked out on ext-alert under its local controls account home directory.
I removed the old ext_alert from h1fescript0's monit. The new service is running under systemd control on ext-alert. To distinguish between the two systems, we have replaced the underscore with a dash in the new system name. I'll update all the documentation, and provide operator instructions on monitoring the system.
Attached MEDM show a Fermi GRB event at Mar 30 2019 01:55:36 PDT this morning.
TITLE: 03/30 Day Shift: 15:00-23:00 UTC (08:00-16:00 PST), all times posted in UTC
STATE of H1: Observing at 109Mpc
OUTGOING OPERATOR: Travis
CURRENT ENVIRONMENT:
Wind: 1mph Gusts, 0mph 5min avg
Primary useism: 0.02 μm/s
Secondary useism: 0.17 μm/s
QUICK SUMMARY: ~45h lock. PEMEX & EY excitations according to the CDS overview, though the DIAG_EXC isn't noticing them and they are not ignored in the code...
Though the PEM excitation channels are open, their value is 0. I will clear these next lock loss.
TITLE: 03/30 Owl Shift: 07:00-15:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Observing at 108Mpc
INCOMING OPERATOR: TJ
SHIFT SUMMARY: Locked for the entire shift, but out of Observe for ~1.5 hours due to EQ. Lock is now ~45 hours long according to Summary Pages.
LOG: See previous aLogs.
CP2's liquid nitrogen storage tank (tank 8514376) may go into alarm sometime before its scheduled delivery on Tuesday. I've forgotten the current alarm level setpoint (had been <20% in the past). This is anticipated and as the result of missed deliveries due to "snowmageddon".
NO ACTION is required by the operators should this go into alarm.
Cell phone alarms will activate when CP2 Dewar level goes below 15%. Since it is already down to 19% I've bypassed this alarm through to Tuesday afternoon.
Bypass will expire:
Tue Apr 2 14:07:35 PDT 2019
For channel(s):
H0:VAC-LX_CP2_LT155_DEWAR_LEVEL_PCT
Returned all setting to pre-EQ values. The SEI Guardians for a couple of platforms reported EZCA connection errors briefly, but cleared themselves before I could get snapshots. No other issues with returning to Observe.
Mag 6.4 (or 6.1 depending on which screen you believe) EQ in Papua New Guinea. Jim's EQ response predictor (see screenshot) suggested taking the SEI Config to Large EQ, which I did. Jim mentioned that earlier in the week some of the Guardian nodes did strange things when switching states, so I took a screenshot of that screen as well which appears to have completed the transition properly. I'll return to the nominal SEI state and Observing when things settle down.
Jenne called to advise me to zero the Master Gain for the ADS loops and to try the new Earthquake state in ISC_LOCK. We are still locked and the ground motion is starting to come down. I'll return the Master Gain to 1 and ISC_LOCK to NLN and get back to Observing shortly.
SDF transient:
2019-03-30_11:16:34.588624Z DIAG_SDF [RUN_TESTS.run] USERMSG 0: DIFFS: susitmy: 1
TITLE: 03/30 Owl Shift: 07:00-15:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Observing at 107Mpc
OUTGOING OPERATOR: Jim
CURRENT ENVIRONMENT:
Wind: 5mph Gusts, 4mph 5min avg
Primary useism: 0.04 μm/s
Secondary useism: 0.17 μm/s
QUICK SUMMARY: Lock is 13:40 long according to the lock clock, and off the chart (+24 hours) on the range FOM.
TITLE: 03/30 Eve Shift: 23:00-07:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Observing at 110Mpc
INCOMING OPERATOR: Travis
SHIFT SUMMARY: Not much to say, quiet shift
LOG:
0:00-0:15 Dropped out of Observe several times by, I think, aborted CAL measurements
3:00 Cheryl driving near VPW
J. Kissel, L. Sun We were all so proud to claim 1%/1deg systematic error in the PCAL 2 DELTAL EXTERNAL transfer function this morning -- see LHO aLOG 48040 -- but this required multiplying the actuator model in the front-end to be scaled by 0.95 (i.e. a 95% decrease). Turns out the problem lies in the detuned sensing function -- namely that a bug in the fitting code falsely reported no residual frequency-dependent systematic error between model and measurement, when -- in fact -- the model was bogus, and there was indeed a substantial residual -- to the tune of ~5% below 20 Hz -- consistent with the features we were trying to remove by scaling the actuator. Since the assumption is that we've falsely put this 5% on the actuator, we are over reporting our sensitivity by 5% below 20 Hz. Since this is so low in frequency, it shouldn't have much impact on the BNS range. However, there must be good agreement between measurement and model (i.e. frequency dependent errors of all functions must be below ~1%) in order for the time dependent correction factor calculations to work, because the calibration lines at all frequencies (but especially those at 20 Hz) are multiplied by loop correction factors produced by the model. I'll explain in more detail below, but the conclusion is that - In order to obtain better agreement, we must better measure the interferometer detuning [namely, measure below 5 Hz], in order to better constrain the MCMC fitting of the spring frequency and Q. - Once we're happy with that, we must update the CAL-CS inverse sensing function filter (again). - Then we can re-verify that - we don't need any fudged factor of 0.95 on the CAL-CS implementation of the actuator model, - the PCAL2DELTAL transfer function is flat, and - the CAL-CS implementation of the model matches the model itself with good fidelity We have received the OK from Keita, the run manager to execute this plan tomorrow (Saturday, 2019-03-30) starting around 11:00a [assuming the IFO is up and running]. DETAILS A few days ago, the primary sensing function code in the pyDARM infrastructure, /ligo/svncommon/CalSVN/aligocalibration/trunk/Common/pyDARM/src/sensing.py was modified to be able to handle a pro-spring detuning which now both L1 and H1 see (e.g. LHO aLOG 47941 or LHO aLOG 42791; of course, L1 sees it to a lesser extent because H1's point absorbers). H1's spring is rather prominent in our current configuration given our choice of SR3 heating (see e.g. LHO aLOG 47604), only decided upon last week for improved noise. This core code is one file with 1000+ lines composed of many functions including, (1) process the collection of sensing function measurements, (2) uses an MCMC algorithm to fit that measurement, (3) generates a model of the sensing function based on the most likely, maximum a posteriori (MAP), values of the MCMC fit, plus a foton design string to invert it (4) divides that new fit model by measurement to to produce an estimate of residual frequency-dependent systematic error, and (5) uses a Gaussian process algorithm to fit the remaining frequency-dependent systematic error to a posterior distribution of frequency-dependent functions. It also has functions that are mere tools like, (A) create the optical plant frequency response of the now-standard detuned-FPDRMI given input of the cavity pole frequency, and detuned optical spring frequency and Q (B) assembling all the details of the full sensing function on to an input optical plant frequency response (C) creating an estimate of the uncertainty given an input coherence (D) creating a Gaussian probability density function from input data vector, with mean, and 1-sigma uncertainty (E) creating a optical plant seed function/equation over which to run the MCMC fit from input parameters but it was (A) and (E) that were modified, and for-better-or-worse, the equations to create the optical response for these two functions were not common to both. In order to invoke the new functionality model of a pro-spring (instead of an anti-spring which had only ever been seen at H1 during O1 and O2), one set the parameter for the spring frequency to be imaginary inside the collection of pyDARM model parameters, e.g. /ligo/svncommon/CalSVN/aligocalibration/trunk/Runs/O3/H1/params/modelparams_H1_20190328.py However, in the modification of (A), the code checked if the input argument was imaginary, and then if so applied the following math for the input arguments detuneSpringFreq (now imaginary) and detuneSpringQ: Line 62: detuneFunction = signal.TransferFunction([1,0,0],[1,2.0*np.pi*detuneSpringFreq/detuneSpringQ,-(2.0*np.pi*detuneSpringFreq)**2]) We discovered this afternoon, that with an imaginary input argument for the spring frequency, this math produces a bogus optical spring response only after comparing the CAL-CS implementation of a spring with fs = 5.620 Hz and Q = 5.201 -- which had been entered into foton directly from the MAP of the MCMC by humans (LHO aLOG 47941), bypassing the (currently broken for pro-spring) foton string from function (3) -- against the pyDARM produced optical plant. This comparison is the first .pdf attachment, 2019-03-29_H1_C_pyDARM_vs_CALCS_beforemodifiedprospring.pdf. For a comparison of what the correct math should be for a pro spring, see the second .pdf attachment, 2019-03-29_bogus_spring_demo.pdf. This bug was fixed by changing the same line of code to be Line 62: detuneFunction = signal.TransferFunction([1,0,0],[1,2.0*np.pi*abs(detuneSpringFreq)/detuneSpringQ,(2.0*np.pi*abs(detuneSpringFreq))**2]) With this fix, the pyDARM code produced a virtually identical response to foton, as expected The new comparison is shown in the third .pdf attachment, 2019-03-29_H1_C_pyDARM_vs_CALCS_modifiedprospring.pdf But -- what does this all mean for the systematic error in the sensing function? Well, after quadruple checking, it turns out function that creates the MCMC fit seeding function (E), was split into two functions (one for pro- and one for anti- spring) and insensitive to this bug. The pro-spring function provided the correct pro-spring formula, Line 739: est_opt_response_tf = mu[0] / (1.0 + 1j*xdata/mu[1]) * (xdata**2/(xdata**2 - mu[2]**2 + 1*xdata*mu[2]*mu[3])) * np.exp( -2.0*np.pi*1j * mu[4]*1.0e-6 * xdata ) with an inputs of a frequency vector xdata and five-element vector of mean value priors, mu = [optical gain, cavity pole frequency, spring frequency, spring Q, and a delay], given that the values for the mu are hard-coded on line 786 (yuck!). The hard-coded prior for the spring frequency is a real, positive 5.0 Hz, i.e. expecting an anti-spring. That meant that the updated MAP values for these parameters from the MCMC fit function, (2), were "right" but the function used to create a model response from those values, i.e. (A), was producing the wrong frequency response that does not accurately represent the MAP values. *That* meant that when this MAP updated model was divided from the measurement, in (4), the amount of residual frequency-dependent systematic error was wrong, and coincidentally small. To see this, look at the fourth and fifth .pdf attachments side-be-side which compare MCMC fit results against model, - before the bug fix 2019-03-29_H1_sensingFunction_beforemodifiedprospring.pdf - after the bug fix 2019-03-29_H1_sensingFunction_modifiedprospring.pdf Since we now are using the fixed, committed function (A), the latter is now physically correct, but sadly, it (finally truthfully) reveals that the MCMC fit to the detuning is actually "quite" poor. In other words, there is not enough low-frequency data to properly constrain the MCMC fit. The right panels of the 3rd page, or all of the fourth page show the true systematic error between model and measurement -- and this true systematic error is consistent with a 5% error [1.05] around 20 Hz, which is consistent with having to scale the actuator down by 5%. You'll notice that in today's broadband and sweep PCAL2DARM transfer functions, well-below the sensor / actuator cross-over frequency, and especially below 20 Hz, the transfer function is asymptoting to 0.95 -- exactly the "correction" that I put in. Using some toy math, you can see this immediately assuming that we don't model the sensing measurement, C_meas, well, then we could fake the answer at the cross-over frequency and above by adjusting the gain on A: dL_real = (1/C_meas) d_err + A_meas * d_ctrl dL_real = (1/C_model) (C_model/C_meas) * d_err + A_meas * d_ctrl (C_meas/C_model) dL_real = (1/C_model) * d_err + (C_meas/C_model) * A_meas * d_ctrl *phew* What an insidious bug!! So, again, we must measure *more* of the detuned spring shape (i.e. to lower frequency) in order to nail down the frequency and Q -- which will help nail down the optical gain, and get everything right in the end. We'll execute the above plan tomorrow.
In an attempt to understand why LLO's kappa calculations are all working nicely, but ours are not, I started looking at the differences between their line demodulation and ours.
First, a change that we are keeping: When looking at these, I found that the local oscillator for Pcal line 2's demodulation had an amplitude of 0, so I'm not sure what it was actually demodulating. The LO wasn't anything other than digital zeros. Sadly, fixing this didn't fix our kappas. Not so surprising, since this is the line for the kappa_c, which is the one kappa that (somehow, even with this problem) was getting the right answer.
A set of changes that we've already reverted: At LLO, they use singain and cosgain values of 1 for their local oscillator demodulators, whereas here at LHO we use values that match those of our actual oscillators. That doesn't seem to have any effect on our kappa calculations (which was surprising actually - why didn't the output values change if the demod outputs were different by, in some cases, several orders of magnitude?). Anyhow, all of these were reverted.
We were out of Observation for a few minutes for this test.
I'm still confused why our kappa calculations aren't working - everything seems like it should be fine.
The wind fence sensors in End X are currently offline. This was found by looking at the H1:SYS-ETHERCAT_X1DEVICE0_ERROR_FLAG. The [CXY]1DEVICE[0-6] channels are monitoring the Beckhoff hardware and typically this would indicate a serious error.
Dave and Jonathan,
The external alert system is not connecting to GraceDb. Looking back at the logs it has not been working since 28 Feb 2019. Queries to gracedb are failing with SSL errors.
In tracking this down we moved the external alerts to a new Debian 9 vm (ext-alert.cds.ligo-wa.caltech.edu) from h1fescript0. This was to put it under better configuration management and to get to a newer set of python and ssl libraries. This has not solved the problem. We are in communication with the GraceDb developers.
I heard back from the GraceDb developers.
When GraceDb moved to being hosted on AWS the load balancer that put in front of the system was changed. The new proxy could not understand impersonation proxy certificates. Simply removing a -p option in the call to ligo-proxy-init when generating the x509 certificate used to authenticate to GraceDb.