Jonathan, Dave:
In the past week we have had several occasions of DAQ-CRC error counts, primarily on h1seiex and also on h1susex. Here is what we have discovered so far:
All CRC errors are accurately reported in the DAQ data concentrator's log file
Errors appear at random times, and at random times within the second
When a front end has errors, for each model, its CRC counter increments by the same number every time
The number the CRC increments by is not the same across models. If it differs, it increments over the model position in the rtsystab file (e.g. IOP < USER1 < USER2)
The error, as seen by the data concentrator, is that the model's gps time is one second in the past
If, during the period of gps mismatch, a second boundary is crossed, both dcu_gps and gps increment within the same 16Hz cycle
The data concentrator marks the DCU DAQ-STATUS internally as 0x4000 (timing error).
This error propagates to the frame writer, which marks all channels from the model with this status
The error does not always show up on the EPICS STATUS PV (perhaps gets reset to zero before being reported?)
h1seiex:
number of CRC events since 11jun2019: 10
model order in mx_stream (num crc_errors each time): h1iopseiex(7), h1hpietmx(8), h1isietmx(8)
Error sequence: (hpi, isi) + 7*(iop, hpi, isi)
h1susex:
number of CRC events since 11jun2019: 2
model order in mx_stream (num crc_errors each time): h1iopsusex(8), h1susetmx(8), h1sustmsx(8), h1susetmxpi(9)
Error sequence: etmxpi + 8(iop, etmx, tmsx, etmxpi)
Swept the EX, EY & just finished off the LVEA at 20:25utc (1:25pm PDT).
(Both End Stations returned to Laser SAFE & LVEA remained Laser HAZARD.)
Yesterday's h1iscey kernel panic was the second in one week, but I seem to remember we went a long period prior to this without these crashes. This prompted me to search the alog and my records to find the freqency of these types of crashes.
| Date | Frontend | global DAQ error |
| Sep 25 | h1lsc0 | yes |
| Sep 28 | h1iscex | yes |
| Oct 06 | h1lsc0 | no |
| Oct 06 | h1susey | no |
| Nov 01 | h1seih16 | no |
| Nov 04 | h1seih16 | yes |
| Jun 11 | h1sush2a | no |
| Jun 17 | h1iscey | yes |
Some observations:
Crashes appear to happen in pairs?
50% of crashes cause DAQ errors on associated systems (same network port on h1dc0), rest are local only.
We went for 7 months before the recent pair of crashes.
At LLO, we compile these into a MASTER Fault ticket - See FRS 12924. I still need to go back and add some older ones.
Managed to finish the OPLEV charge measurements today during the maintenance downtime, for both the ends (ETMX and ETMY). The results are attached below. The effective bias voltage for both ETMX and ETMY this week has been mostly below 40 V (except for ETMY yaw which is now 50V). The long term trend is on the rise for some of the quadrant, however we still don't need to flip the sign yet.
After the measurements were taken, all the values were restored to it's orignal and SDF differences have been removed (expect for optical align offset for pitch and yaw for both ETMX and ETMY which is still flagged, maybe operators can accept the difference if required).
At the start, I clarify that these above measurements that Rahul regularly takes are measures of *absolute* *angular* actuation strength of each ETM ESD system, in terms of effective bias voltage in Volts on the ESD bias electrode (i.e. an imagined *extra* bias voltage on top of the regularly *requested* bias voltage [think ~100s of Volts] that would bring this absolute angular actuation strength to *zero* when the requested voltage is set to zero).
I re-clarify all this terminology because conversations are happening about whether we'll need to
- launch a Test Mass Discharge System campaign any time soon,
- begin alternating the requested bias voltage
- do nothing
during O3.
I answer the question with additional plots of another, different metric of the charge accumulated: the *change* in *relative* *longitudinal* actuation strength for the ESD system on ETMX (alone) between
- April 18th 2019 (during O3, after we think we finally got the calibration tracking system right),
- May 18th 2019, and
- June 18th 2019.
Remember that the *absolute* *longitudinal* actuation strength for the ESD was measured to be 4.739e-12 N/ct by the calibration group on Mar 28th 2019 (see LHO aLOG 47941), and this relative longitudinal actuation strength is with respect to that. Note, you can find these plots on any day, readily made under the "CAL > Time Varying Factors" tab of the individual IFO summary pages, and these are plots of the real part of the "\kappa_TST" factor (e.g. ).
The message: the *longitudinal* actuation strength for ETMX has not changed since Mar 28th 2019 (about 4 months) by more than 1%. (Probably less than that, but difficult to quantify rigorously from the plot alone since there are so few tick marks on the y-axis.)
Recall, in O1 and O2, we were seeing changes on the order of 0.5% *per week* in this relative longitudinal actuation strength (e.g. LHO aLOG 24241) -- which motivated us to develop a campaign of requested bias voltage sign flipping (see LHO aLOG 33985 and references there-in).
Also recall that, during O1 and O2, we were not only concerned about relative actuation strength changing -- which impacts the calibration of the detector -- but also the noise contributions of charge to the detector sensitivity. This was identified by shaking BSC-ISIs and measuring the direct coupling the sensitivity in the 10-50 Hz region (when one would naively expect any nominal such excitation to be filtered aggressively by the QUAD suspension with a slope of 1/f^8). We indeed found such coupling -- see e.g. (LHO aLOG 37752).
However, finally recall that, we then successfully deployed the Test Mass Discharge System at each end station for at LHO (ETMX LHO aLOG 38457, ETMY LHO aLOG 38524), and clearly successfully reduced the effect on noise on ETMX (LHO aLOG 38507) and successfully reduced that same test masses estimate of effective bias voltage (as measured by optical levers LHO aLOG 38604). (At the time, we were using ETMY as our DARM actuator, and it's absolute longitudinal actuation strength was remeasured, thus resetting the relative longitudinal actuation strength back to 0%).
My suspicion is that now that we've mitigated sources of free ions, e.g.
- Moving ion pumps 250 m away from the test mass
- Installed chevron baffles around those ion pump intakes
- developed a program of discharging the test masses upon exit from the chamber,
that we've successfully created a system which no longer charges up, even though we're still regularly operating with the ETMX ESD system in use at 70% duty cycle and a ~400 [V] requested bias voltage.
We have not yet remeasured the BSC-ISI to DARM coupling in O3, so we don't yet know of any direct noise coupling, so this assessment is positive, but a bit incomplete until those measurements are redone. However, for now, I think we can take option 3:
- We do not need to do or change anything thus far in O3; charge is not a problem.
We will, of course, remain vigilant and continue regular measurements, as we all know that as soon as we stop measuring we'll have another 6.7 Mag earthquake in Montana.
I did a little bit of PRMI ASC this morning while we were unlocked because of the earthquake but before maintence day was in full swing.
INIT_ALIGN
As Patrick is bringing back the IFO from maintenance, the PRMI locks are have very low POP18 and POP90 values, and it looks like we might be touching the low side of the gain bubble, because things keep oscillating, then we lose the PRMI lock before the PRMI ASC can come on.
I tried by-hand changing the CHECK_MICH_FRINGES state to a mich bright state (just hand flipped the sign of the gain), and then tried the initial alignment mich bright ADS settings, which dither the BS, demodulate AS_DC, and actuate on the BS. I'm not sure why, but this didn't seem to bring the BS to a better alignment. I only tried twice since I don't want to delay the IFO recovery by much, but I think that we should someday investigate this as a follow on to the PRMI work that Sheila has done today.
A few more guardian changes:
(Almost) One click initial alignment:
PRMI changes in ISC_LOCK:
I reset both PSL power watchdogs at 17:19 UTC (10:19 PDT). This closes FAMIS 10715.
Also, while resetting the watchdogs I noticed the Inner Loop ISS diffracted % was down around 1%. This should be between 2%-3%, so I adjusted the ISS RefSignal from -2.07 V to -2.05 V; this brought the diffracted % to ~2.0%. This change was accepted in SDF.
All baluns at the corner station have been replaced with improved v3 modules, ALOG-48165. Twelve v1 baluns remain at the end stations and will be replaced next maintenance Tuesday.
TITLE: 06/18 Day Shift: 15:00-23:00 UTC (08:00-16:00 PST), all times posted in UTC
STATE of H1: Earthquake
OUTGOING OPERATOR: TJ
CURRENT ENVIRONMENT:
Wind: 17mph Gusts, 10mph 5min avg
Primary useism: 0.42 μm/s
Secondary useism: 0.14 μm/s
QUICK SUMMARY:
SEI_CONF set to LARGE_EQ_NOBRSXY for earthquake in Japan. Start of maintenance. Contractor has started on LN2 Dewar fill lines (WP 8243).
TITLE: 06/18 Owl Shift: 07:00-15:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Earthquake
INCOMING OPERATOR: Patrick
SHIFT SUMMARY: Quite shift until a 6.5 earthquake from Japan. Maintenance has started.
LOG:
6.5mag from Japan. Showed up much earlier than seismon predicted, and much more suddenly than I could react to.
I have switched to LARGE_EQ_NOBRSXY and so far nothing has tripped.
The EQ finally showed up on the USGS webpage at 1344 UTC
Interesting. At Livingston we were a bit farther away, so the P-wave (which did not knock us out) was reported by seismon and the S-wave (which did knock us out of lock) hit us almost exactly when seismon said it would.
8.5 hrs lock, wind has died down. Smooth sailing.
At 01:27 (18:27) got a CRC error due to problem with H1IOPISCEY. Called Dave B., see aLOG 50009. Lost lock at 01:53 (18:53) as expected when Dave rebooted H1ISCEY.
Reset End-Y ISI and HEPI tripped WDs.
At 02:01 (19:01) started relocking. Initially had some issues getting Y-Arm Green to lock. At the time there were several wind gusts in the mid 30mph range. After things settled a bit were able to reliably lock both arms green. Had to tweak the BS and PRM to get through DRMI_1F. After this no additional adjustments were needed.
At 03:02 (20:02) and 03:23 (20:23) lost lock at ENG_SOFT_LOOPS. Georgia is looking into these lock losses.
Have been seeing 30 plus mph gusts at End-X, which is not helping the relocking.
We lost lock engaging the soft loops a few times this evening.
It looked like the corner ASC could not keep up when we switched on the arm dither loops, symptom: power recycling gain goes up, rf18 buildup drops, we lose lock.
I removed 10dB of gain from the arm dither loops (pit4, pit5, yaw4, yaw5) when they first come on (set in PREP_ASC_FOR_FULL_IFO), and added the 10dB back in later in ENGAGE SOFT LOOPS. This worked by hand, so I put it in the guardian.
Dave was thinking we could restart this ISCEY computer without affecting the SEI or the SUS platforms and we all agreed until we realized we shouldn't have. There was no impact on the models of the SEI or the SUS but, ALS is bleeding the Tidal Drive from the ETM to the HEPI. When the ALS came back on with zero Tidal Drive, the HEPI was jerked back to zero offset. The HEPI tripped on this 6 seconds before the ISI tripped when the T240 tilted off or got slung sideways by the HEPI. Attached is 4 hours showing the tidal drive on HEPI coming from the ALS (lower left) going flat and then jerking this way and that on its way to zero when the other channels, the ISI's T240s, start railing.
I find interesting that the T240 signals all go flat when the ALS does with the Kernel panic but maybe that would be expected too.
We'll endeavor to implement some procedure to hold this output on the HEPI and then bleed it off at an acceptable rate.
The general cores on h1iscey have locked up. The models on this front end continue to run (pemey, iscey, alsey, caley) but their EPICS and DAQ data is invalid.
All front ends which share h1iscey's port on h1dc0 have invalid data, which should clear once mx_stream is going again on h1iscey.
I am starting the reboot process.
I have recovered h1iscey by remotely resetting it via its IPMI port. Procedure was:
Once the mx_stream was restarted on h1iscey, the DAQ errrors on the other system cleared.
Handing over to Jeff B for IFO recovery.
Details, taking h1iscey out of Dolphin fabric and verification:
david.barker@zotws6: ./dolphin_switch_port_status.sh h1iscey
h1iscey Switch_IP 10.101.0.94 Switch_port 3 is ENABLED
david.barker@zotws6: dolphin_disable_switch_port h1iscey
Disabling h1iscey Switch_IP 10.101.0.94 Switch_port 3
david.barker@zotws6: ./dolphin_switch_port_status.sh h1iscey
h1iscey Switch_IP 10.101.0.94 Switch_port 3 is DISABLED
Details, remote IPMI chassis reset
david.barker@zotws6: ipmitool -I lan -H 10.99.101.193 -U (user) -P (passwd) chassis status
FRS13078 opened to cover this issue.
M. Pirello, D. Gustafson
We installed new v3 Baluns on the following signals in the CER:
ISC-C3(37-1) 40Mhz TCS AOM Return exchanged v1-117 with v3-S002
ISC-C3(33-4) 158.8Mhz Fiber Beat Note Return exchanged v1-193 with v3-S003
ISC-C3(33-5) 79.4Mhz SQZ VCO Return exchanged v1-151 with v3-S004
ISC-C3(33-6) 203.125MHz SQZ VCXO Return exchanged v1-052 with v3-S005
ISC-C4(39-3) 79.4MHz PSL VCO Return exchanged v1-093 with v3-S010
These are all return signals, the impact should be minimal. Work was completed per WP8150
Balun status can be seen here E1900100.
Continuing the Balun exchange program:
We installed new v3 Baluns on the following signals in the CER:
ISC-C4(39-1) 79.4MHz "ALS DIFF VCO Return" exchanged v1-102 with v3-S006
ISC-C4(39-2) 79.4MHz "ALS COMM VCO Return" exchanged v1-094 with v3-S068
ISC-C4(26-2) 45.5MHz "PEM Readback for Antenna Demod" exchanged v1-055 with v3-S009
ISC-C4(19-8) 9.1MHz "PEM Readback for Antenna Demod" exchanged v1-147 with v3-S116
Work was completed per WP8190. WP was modified to reduce impact to the IFO during this weeks maintenance, we did not touch squeezer.
Continuing the Balun exchange Program:
We installed new V3 baluns on the following 45MHz signals in the LVEA:
ISC-R1 (41-6) 45MHz Auxiliary Modulation (EOM Driver) exchanged v2-DG146 with v3-S152
ISC-R2 (41-2) 45MHz Distribution exchanged v2-DG083 with v3-S094
ISC-R3 (41-3) 45MHz Distribution exchanged v1-145 with v3-S117
Work was completed per WP8200, balun status can be seen here E1900100.
I have attached Insertion Loss and Leakage scans from the old baluns as well as the new ones.
M. Pirello, D. Gustafson
Continuing the Balun exchange Program:
We installed new v3 baluns on the following signals at the PSL racks and at HAM6:
PSL-R2 (18-2) ISS AOM 80MHz exchanged DG-104 with v3-S069
ISC-R3 (41-2) Distribution 4th Harmonic exchanged v1-053 with v3-S160
ISC-R3 (41-4) Distribution 8th Harmonic exchanged v1-091 with v3-S040
ISC-R3 (39-4) Distribution 42.2MHz exchanged v1-125 with v3-S060
Work was completed per WP8215, balun status is in the same place it was last week, E1900100.
Attached is a comparision between the balun modified by hand (DG-104), and the new v3 balun which replaced it.
M. Pirello, P. King, D. Gustafson
Continuing the Balun exchange Program:
We installed new v3 baluns on the following signals at the PSL racks:
PSL-R2(17-1) FSS Modulation exchanged v1-043 with v3-S157
PSL-R2(17-2) FSS exchanged v1-014 with v3-S093
PSL-R2(17-4) PMC Modulation exchanged v1-021 with v3-S115
PSL-R2(17-5) PMC exchanged v1-050 with v3-S156
PSL-R2(17-6) Injection Locking exchanged v1-084 with v3-S095
ISC-R1(41-2) Distribution JAC Future exchanged v1-159 with v3-S092
** On the last one, the labels on the feed through may be more correct than the DCC. I followed the cable path to the ALS VCO for the PSL.
If it is the ALS VCO, Peter told me that it is possibly 80MHz and in this case the difference between V1 and V3 is about 0.75dB more power, and no measurable difference in phase. The PSL signals being under 45Mhz should be less than 1 degree difference in phase, and less than 0.25dB difference in power.
Work was completed per WP8208, balun status is in the same place it was last week, E1900100.
I have attached insertion loss and leakage scans of v1 vs v3 baluns.
We checked the ISC-R1 (41-2) Balun and determined this signal is the REFL_B demod, 9MHz. There should be very little phase & power difference between V1 and V3 baluns at this frequency.
M. Pirello, D. Gustafson
Continuing the Balun exchange Program:
We installed new v3 baluns on the following signals at the end stations.
EX-ISC-C1 (41-5) Return PLL Beat Note 39.5MHz exchanged v1-176 with v3-S050
EX-ISC-C1 (41-6) Return ALS Laser VCO 79.4Mhz exchanged v1-049 with v3-S042
EX-ISC-R1 (41-1) ALS Laser VCO 71MHz exchanged v1-189 with v3-S114
EX-ISC-C1 (41-5) Return PLL Beat Note 39.5MHz exchanged v1-137 with v3-S041
EX-ISC-C1 (41-6) Return ALS Laser VCO 79.4Mhz exchanged v1-199 with v3-S155
EX-ISC-R1 (41-1) ALS Laser VCO 71MHz exchanged v1-039 with v3-S059
Work was completed per WP8222, balun status is in the same place it was last week, E1900100.
M. Pirello, D. Gustafson
Continuing the Balun exchange Program:
We installed new v3 baluns on the following signals at the Squeezer
ISC-R3 (39-1) 80MHz SQZ EOM (OPO) exchanged v1-178 with v3-S082
ISC-R3 (39-2) 35.5MHz SQZ EOM (SHG) exchanged v1-060 with v3-S089
ISC-R3 (39-3) 200MHz SQZ EOM (CLF) exchanged v1-166 with v3-S031
SQZ-R1 (41-1) 80MHz OPO Demodulation exchanged v-122 with v3-S084
SQZ-R1 (41-2) 35.5MHz SHG Demodulation exchanged v1-020 with v3-S033
SQZ-R1 (39-3) 71MHz SQZ VCO Laser Locking exchanged v1-190 with v3-S083
Work was completed per WP8230, balun status is in the same place it was last week, E1900100.
** The 200MHz SQZ EOM CLF signal may require a slight phase change, about 8 degrees difference and 2dB more power with the V3.
M. Pirello, D. Gustafson
Continuing the Balun exchange Program:
We installed new v3 baluns on the following signals at the Squeezer and the PSL racks:
SQZ-R1 (41-3) 3.125MHz SQZ Angle Demodulation exchanged v1-171 with v3-S108
SQZ-R1 (41-4) 3.125MHz SQZ Angle Demodulation exchanged v1-145 with v3-S109
SQZ-R1 (41-5) 6.25MHz CLF Demodulation exchanged v1-183 with v3-S085
ISC-R1 (41-1) 71MHz Distribution exchanged DG-02 with v3-S086
ISC-R1 (41-3) 15th Harmonic Modulation exchanged DG-01 with v3-S058
TCS-MEZ (41-1) TCS exchanged v1-153 with v3-S079
Work was completed per WP8240, balun status is in the same place it was last week, E1900100.
M. Pirello, D. Gustafson
Continuing the Balun exchange Program:
We installed new v3 baluns on the following signals at the field racks near the PSL:
ISC-R2 (41-1) 9MHz Distribution exchanged v1-047 with v3 S019
ISC-R2 (41-3) 2nd Harmnonic Distribution exchanged v1-098 with v3-S019
ISC-R2 (41-4) 3rd Harmonic Distribution exchanged v1-157 with v3-S015
ISC-R2 (41-5) 10th Harmonic Distribution exchanged v1-179 with v3-S016
ISC-R2 (41-6) 15th Harmonic Distribution exchanged v1-074 with v3-S132
ISC-R1 (41-4) 24MHz MC Distribution exchanged v1-040 with v3-S017
ISC-R1 (41-5) 9MHz Main Modulation exchanged v1-033 with v3-S018
Work was completed per WP8244, balun status located at this link: E1900100. This concludes all "known" balun work at the corner station. We are sprinting to next week where we intend to replace the remaining twelve v1 baluns at the end stations.
M. Pirello, D. Gustafson
Continuing the Balun exchange Program, we installed new v3 baluns on the following signals at both end stations:
EX
ISC-R1 (41-2) 24.4MHz Modulation exchanged v1-022 with v3-S074
ISC-R1 (41-3) 24.4MHz Demodulation exchanged v1-184 with v3-S131
ISC-R1 (41-4) Not Connected exchanged v1-064 with v3-S078
ISC-R1 (39-1) 24.4 WFS A Demod exchanged v1-077 with v3-119
ISC-R1 (39-2) 24.4 WFS B Demode exchanged v1-026 with v3-S076
ISC-R1 (39-3) 71 CPS Timing Fanout exchanged v1-059 with v3-S020
EY
ISC-R1 (41-2) 24.4MHz Modulation exchanged v1-127 with v3-S073
ISC-R1 (41-3) 24.4MHz Demodulation exchanged v1-118 with v3-S011
ISC-R1 (41-4) Not Connected exchanged v1-013 with v3-S012
ISC-R1 (39-1) 24.4 WFS A Demod exchanged v1-126 with v3-S075
ISC-R1 (39-2) 24.4 WFS B Demod exchanged v1-161 with v3-S013
ISC-R1 (39-3) 71 CPS Timing Fanout exchanged v1-191 with v3-S135
Work was completed per WP8255, balun status located at this link: E1900100.
** In the process of walking through the LVEA we found 4 more Baluns hidden among the TCS and SUS racks which need upgrading. We were able to upgarde one of these with this Tuesday, and will finish the remainder next maintenance day.
SUS-R3 (40-1) CPS HAM 71MHz exchanged v1-011 with v3-S113
This additional work was started per WP8261
M. Pirello, H. Radkins
The final three baluns were exchanged at the corner this morning for a total of 64 baluns replaced over 14 weeks.
TCS-R2 (41-1) 71Mhz CPS Timing exchanged v1-088 with v3-S144
TCS-R2 (41-2) N/A Empty signal with Balun exchanged v1-081 with v3-S001
TCS-R1 (41-6) TCS Signal from Mechanical Room exhcnaged v1-143 with v3-S045
All work finished per WP8261. This completes work done per FRS9794 and ECR-E1700404 at LHO.
Attached plot of full data at the time of this morning's h1susex CRC glitch shows that for 8 cycles (0.5 seconds) the DAQ data is repeated. It is easy to see in a slowly varying channel (lower plot) and more hidden in the upper plot.
We reviewed the code and understand the internal mechanism of the error. There are several possible causes:
the front end models are slow in getting data into shared memory*
mx_stream is slow in reading the DAQ data from shared_memory
mx_stream is slow in processing the data (it calculates the data crc)
mx_stream is slow in sending the data
the network is slow in getting the data to the data_concentrator
* this is highly unlikely, however the mx_stream is triggered by the IOP being ready, and of the 4 computers at EX the two with problems have iop watchdogs.
Outstanding questions are;
Why is this happening only at EX?
Why did it only start last week?
Why only in h1seiex (10) and h1susex (2)?
We are reviewing our options for mitigation and/or further diagnostics.
Jonathan, Greg, Dave:
We have verified that both the full frames and the HofT frames are correctly marking the DAQ status of all impacted channels as bad (0x4000 = timing error).
Opened FRS13093 to cover this