(& a large dash of microseism)
Have been receiving Alarm Handler alarms for MX low temperatures; one sensor is at/below 60°F. Attached is the last 2+ months of temperatures at MX.
Seeing gusts over 50mph here at the Corner Station.
Also can make out the tumbleweeds beginning their reclammation of the X-arm's X1. Sorry, Tyler/Chris/Scott.
Needless to say, H1 is randomly losing lock all over the place depending on when a building is slammed by wind. Alignment currently looks good with us getting as far as DARM OFFSET once, but the winds are all over the place. Maybe I'll take a break and try an alignment, but really I need the winds to take a break. :-/
Took SEI_CONF to WINDY (from MICROSEISM_WINDY).
Took SEI_DIFF to DOWN, but then went back to FULL.
10:00 UTC Winds are easily averaging 25mph with gusts up to 50mph, and speeds are currently trending up. H1 alignment is good with DRMI locking.
Going to HOLD at LOCKING ARMS GREEN until winds die down.
TITLE: 01/11 Eve Shift: 00:00-08:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Wind
INCOMING OPERATOR: Corey
SHIFT SUMMARY: lockloss due to wind, relocked, lockloss due to wind and useism, Corey's relocking
LOG:
During relocking I used SEI in EARTHQUAKE mode to get the arms to lock and make it through CHECK_IR
I was aware that the ground motion was effecting the H1 range (wind and useism), the SEI guildance was for WINDY. From my earlier experience I was not sure that going to MORE_WINDY would save the lock, and actually thought that changing SEI in any way, in these conditions, would likely break the lock, so I let it play out in WINDY.
VIOLIN Mode gain changes in lscparams, saved and loaded:
TITLE: 01/11 Owl Shift: 08:00-16:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Wind
OUTGOING OPERATOR: Cheryl
CURRENT ENVIRONMENT:
SEI_CONF state: USEISM_WINDY
Wind: 34mph Gusts, 28mph 5min avg
Primary useism: 0.07 μm/s
Secondary useism: 0.67 μm/s
Winds trending up to avg speds of 25mph (& gusts going 45mph). Microseism has quickly (over last 6hrs) increased to above the 90th percentile.
QUICK SUMMARY:
STATE of H1: Observing at 116Mpc
CURRENT ENVIRONMENT:
SEI_CONF state: WINDY
Wind: 17mph Gusts, 13mph 5min avg, 3 hours ago, at 00:38UTC, wind spikes close to 60mph
Primary useism: 0.07 μm/s
Secondary useism: 0.45 μm/s
Attached is a one day plot, ending at the current time, which shows that we've had a number of 50+mph winds. Lockloss at 00:38UTC, 11 Jan 2020, caused by a spike above 50mph. It looks like the lockloss 8.5 hours ago, at 19:23 UTC, 10 Jan 2020, could also be due to winds.
TITLE: 01/11 Eve Shift: 00:00-08:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Observing at 118Mpc
OUTGOING OPERATOR: TJ
CURRENT ENVIRONMENT:
SEI_CONF state: WINDY
Wind: 26mph Gusts, 20mph 5min avg
Primary useism: 0.08 μm/s
Secondary useism: 0.40 μm/s
QUICK SUMMARY: locked in Observe
[Jason, Niko, TJ, Jenne]
After a lockloss from RESONANCE that tripped the ISI-BS_STAGE_2 watchdog, FSS was having problems autolocking. Jason and TJ managed to get it re-locked, but later in locking POPAIR_B_RF18 would drop at CARM_TO_ANALOG and we would lose lock shortly afterward. Jenne noticed that the FSS common gain had been changed, and since this value is not hardcoded in GUARDIAN (because it gets changed whenever the FSS is adjusted by the PSL crew), it remained at the wrong value after we would lose lock. We're not sure how/why it got changed, perhaps if the FSS GUARDIAN reset partway through its common gain scan...(to be continued)
The FSS common gain was changed back to 20dB from -3dB, and we have reached NLN.
[Jenne, Rahul]
With LLO's initial success using the new R0 tracking (LLO alog 50897), we have also put in infrastructure for us to try here at LHO. This became a somewhat bigger project than we were intending, since there were discrepancies in what we are running versus what is in the SVN for the common library parts. This has been an ongoing conversation, but today we separated the ETM and ITM Quad_Master library parts by copying them into the h1-specific folder so that we didn't pull in LLO's changes that they've made for 3.3 Hz damping. Then we made the changes to implement the R0 tracking.
On Tuesday, we'll do the make-installs, and reboot all 4 quad models. As LLO did, we must ensure that the gains to these filter banks are OFF before we take the suspension out of SAFE after the model is restarted, so we aren't sending the raw L2 witness channels straight out to the R0 actuators.
LLO has shared with us the filters that they used in LLO alog 50889, so we'll probably use those as a first try for implementing the R0 tracking. We might need to check if the ITM filters can be the same as the ETM filters, since there are some differences between the ETMs and ITMs.
Then, perhaps with injections (so that we're not reliant on the environment to stay constant), we'd like to check with some on/off tests to see the effect on scattering arches. Probably we'll ask for some commissioning time Wed AM, after we've got these installed on Tuesday.
TITLE: 01/10 Day Shift: 16:00-00:00 UTC (08:00-16:00 PST), all times posted in UTC
STATE of H1: Observing at 115Mpc
INCOMING OPERATOR: Cheryl
SHIFT SUMMARY:
LOG:
h1nds1 restarted at 14:30 this afternoon following a restart this morning. We are investigating this instability and will post details here.
Taking a look at system metrics it looks like the h1nds1 received went under a some short periods of load that exhausted the memory of the system. See the ndscope plot. The two time axis markers are at daqd/nds crashes.
Reading the system dmesg also points out an out of memory condition. Here is a sample from the dmesg output
[6742427.444553] Out of memory: kill process 1350 (nds) score 669513 or a child [6742427.444555] Killed process 1350 (nds) vsz:8034156kB, anon-rss:7857068kB, file-rss:4kB [6742427.458197] nds: page allocation failure. order:0, mode:0x201da [6742427.458199] Pid: 1350, comm: nds Not tainted 2.6.35.3 #5 [6742427.458201] Call Trace: [6742427.458207] [] __alloc_pages_nodemask+0x58a/0x5e3 [6742427.458211] [] alloc_pages_current+0xa2/0xc5 [6742427.458214] [] __page_cache_alloc+0x75/0x7c [6742427.458217] [] __do_page_cache_readahead+0x90/0x1a4 [6742427.458219] [] ondemand_readahead+0x122/0x19c [6742427.458221] [] page_cache_sync_readahead+0x38/0x3a [6742427.458223] [] generic_file_aio_read+0x24d/0x56a [6742427.458226] [] nfs_file_read+0xa9/0xd0 [6742427.458228] [] do_sync_read+0xc6/0x103 [6742427.458230] [] vfs_read+0xa2/0xdf [6742427.458232] [] sys_read+0x45/0x69 [6742427.458235] [] system_call_fastpath+0x16/0x1b
If this is a continued issue then we should switch the control room to reference h1nds2 which is a newer system with more resources.
Attached images show memory available (percent) as reported by h1nds1 around the times of the morning and afternoon crashes. The vertical lines going to 90% denote the time daqd is restarted. These agree with the generation of a new daqd log file. The times are (local PST):
| 10:33 |
| 10:34 |
| 10:49 |
| 14:29 |
| 14:30 |
In both cases the available memory started trending to zero about 13 minutes before the eventual crash. As reported in alog 54406 the 10:49 crash showed no error in the daqd logfile, but dmesg shows a segfault. The afternoon crash gives a dmesg memory error.
The pair or restarts following the memory depletion have the same sequence:
First restart: Retransmissions then packet skip
Second restart: Invalid broadcast received
When the available memory is being depleted, it happens in steps. The width of the steps are roughly 30 seconds, suggesting data requests are being made with that periodicity.
A quick end-of-the-week status update on the seemingly never-ending tumbleweed battle: X arm work is ongoing. Unfortunately, the rain/hail/snow is not doing us any favors with the harvester. At our current pace & barring any significant changes with weather/equipment I would anticipate another week (~5 days) of work to clear this arm. Y arm has 2 rather insignificant blockages which we have yet to tackle but remains impassable. Happy Friday and thanks for everyone's continued patience. Tyler G. Chris S. Scott L.
While RGA manufacturer was on site testing Labview-based RGA software, we noticed a significant AMU 28 peak at EX, towering water signature, suggesting an air leak. The total pressure is 1.9e-9 Torr and the trend is relatively flat, so no urgent concern. I suspect this is due to the failed annulus ion pump on GV20. We may have an inner o-ring leak on GV20 in addition to diffusion through o-rings.
AIP has been off for over a week as EX has been inaccessable due to tumbleweed blockages.
EY scan is also attached for comparison.
The old bash scripts switch_fec_sdf_to_safe and switch_fec_sdf_to_observe have now been replaced with a python program called switch_fec_sdf
usage: switch_fec_sdf [-h] subsys reference
Like its precedessors, for each model name which matches the subsys string, it switches the model's SDF reference to use its safe.snap or OBSERVE.snap file.
The code requires a confirmation, and lists the models which are to be switched.
# switch_fec_sdf h1ioppem observe
Please confirm the switching of the following models to OBSERVE
h1ioppemmx
h1ioppemmy
[y/N]:
The old switch_fec_sdf_to_safe and switch_fec_sdf_to_observe scripts are now wrappers to the new program.
Please email me if you have any problems with the new program.
Seemed very fast, no obvious reasons.