TITLE: 09/11 Day Shift: 15:00-23:00 UTC (08:00-16:00 PST), all times posted in UTC
STATE of H1: Corrective Maintenance
OUTGOING OPERATOR: Corey
CURRENT ENVIRONMENT:
SEI_CONF state: WINDY
Wind: 3mph Gusts, 1mph 5min avg
Primary useism: 0.01 μm/s
Secondary useism: 0.09 μm/s
QUICK SUMMARY:
Sheila is here to asist wit the re-locking efforts following the computer crash.
TITLE: 09/11 Owl Shift: 07:00-15:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Observing at 115Mpc
INCOMING OPERATOR: Ed
SHIFT SUMMARY:
LOG:
Richard, Corey, Dave:
I disabled Dolphin EY port 4, Richard moved the RFM cable from port 4 to port 5 and we now have IPC traffic in both directions to EY.
CDS overview is now all GREEN, looks like our recovery from h1boot1 crash is complete.
I've modified the configuration file on h1boot1 (/diskless/root/etc/ixnodetab) to move h1cdsrfm Y0 to port 5.
h1susey Y0 1
h1seiey Y0 2
h1iscey Y0 3
h1cdsrfm Y0 5 002353
The end-station problem is horribly like the problems we had at LLO last week replacing l1iscex - see LLO log 48295. Again one port on the switch would not respond to port enable/disable commands. As you did, we just moved it to another port. We tried using the 'RESET' command on the web interface (suggested by the vendor previously), but that was insufficient. It may requires a power cycle after the reset, but it is unclear. I am elevating this issue with the vendor to get a resolution.
DQ Shifter: Vaishali Adya
Email: vaishali.adya@anu.edu.au
Mentor : Paul Altin
Fellow(s): Kara Merfeld
Summary (highlights) of the DQ Shift for an aLog:
Full report: https://wiki.ligo.org/DetChar/DataQuality/DQShiftLHO20190902
h1cdsrfm is back up and running, long range dolphin is working except from EY to CS. I suspect a left over from July's problems, I'm working on this.
Also h1edc appears to have lost connection to 90 digital video channels?
Configuration looks fine. Back on 24th July 2019 we moved the long range dolphin EY switch port from 6 to 4 (port 6 was wedged), and I had changed the configuration accordingly. I just did a switch port enable for h1cdsrfm and it correctly enabled port 4 at EY.
Problems with EY Port 4 now?
h1boot1 etc # ./dolphin_ix_remote_enable_bootserver.sh h1cdsrfm
--enable
Enabling switch_no 10.101.0.90, port 5
Complete - /opt/DIS/var/smbus: Read-only file system
not a valid lock (fd=-1)
csr write addr=0x00014050, val=0x20820080 (with ret=0)
--enable
Enabling switch_no 10.101.0.93, port 7
Complete - csr write addr=0x00004050, val=0x20820080 (with ret=0)
--enable
Enabling switch_no 10.101.0.94, port 4
Complete - csr write addr=0x00008050, val=0x20820080 (with ret=0)
IPC MEDM sans EY->CS traffic
Corey, Dave:
Corey power cycled h1susey, h1seiey and h1iscey. This cleared the DAQ errors with EX and corner station ISC. h1susey is running, both SEI and ISC IOP models have a negative IRIG-B excursion which should clear soon. In the mean time their DAQ and IPC status is bad.
Remaining IPC errors are because h1cdsrfm is not responsive. I've taken this out of Dolphin and am remotely reseting it.
After the trip to EY, Dave said I could bring back all the tripped SUS & SEI at EY. Started with ETMy & TMS. Then moved on to HEPI, and then SEI ETMy.
Did notice HPI_ETMY_SC a message for this node (via SEI_CONF) which said "Configuration is not correct, re-request to recover"....there was also some a YELLOW indicator for this node. After re-requesting CONFIG_SC_ON, this Node was good.
This made everything for EY GREEN.
Noticed YELLOW for the B3_SEI & B1_SEISEI guardian nodes. Which showed SEI_ITMx (&y) had notifications for their ISI ITM ST2 guardian nodes for both x & y say "SETPOINT CHANGES. see SPM DIFFS for differences". Not sure what to do about those SPM Diffs. But other than that, all EY SEI guardian nodes are green.
After h1boot1 came back (took a few minutes, and FSCK was needed) the overview MEDM came back to life.
It looks like h1iscey is the only casualty, and it has taken the DAQ data down for all the other front ends which share its h1dc0 port (both ends, corner station ISC and SUSAUXB123).
I'm starting the h1iscey reboot process.
Resetting h1iscey now. Looking through the logs we had an issue with this front end in June of this year.
https://alog.ligo-wa.caltech.edu/aLOG/index.php?callRep=50009
I cannot ssh into any EY Dolphin'ed front end machine (h1susey, h1seiey, h1iscey), even though SUS and SEI EPICS appears operational.
I took h1iscey out of Dolphin and remotely reset it via IPMI. This crashed the models on h1susey and h1seiey. Also, h1iscey was pingable for a short while and then froze out again.
I tried to remotely reset h1susey via IPMI, but its management port is unresponsive.
Corey is now driving to EY to capture the consoles of these three machines and then power cycle them. I'm his safety buddy while he is traveling to and from EY.
Looking around I can ping the front end machines but cannot ping the boot machine, Corey is taking a look at h1boot1's console to see if it needs rebooting.
The DAQ and the slow controls systems EPICS IOCs continue to run, only the front end systems appear frozen, which points to a h1boot1 problem.
Corey reports h1boot1 console has errors, and is unresponsive to keyboard input. We are rebooting h1boot1 via the front panel RESET button.
[Edit: Originally thought this was a Guardian Issue, but we later found out it was due to the h1boot1 computer.]
New behavior to me tonight occurred when attempting an INITIAL ALIGNMENT. Everything was fine up until INIT ALIGN wanted to run the MICH BRIGHT step. This step runs a DOWN command for ALIGN_IFO, but then in an early step of the DOWN code it repeatedly gets an connection error (for PRC1_P). After this, ALIGN_IFO goes into a FAULT state & a continuous loop. Below is the log:
2019-09-11_09:24:03.821637Z ALIGN_IFO REQUEST: MICH_BRIGHT_ALIGN
2019-09-11_09:24:03.822289Z ALIGN_IFO calculating path: DOWN->MICH_BRIGHT_ALIGN
2019-09-11_09:24:03.822719Z ALIGN_IFO new target: PREP_FOR_MICH
2019-09-11_09:24:06.469708Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP1_P => OFF: INPUT
2019-09-11_09:24:06.824697Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP2_P => OFF: INPUT
2019-09-11_09:24:06.830620Z ALIGN_IFO REQUEST: DOWN
2019-09-11_09:24:06.831255Z ALIGN_IFO calculating path: DOWN->DOWN
2019-09-11_09:24:06.831255Z ALIGN_IFO new target: DOWN
2019-09-11_09:24:08.924272Z ALIGN_IFO [DOWN.main] USERMSG 0: EZCA CONNECTION ERROR: Did not observe effect of writing value to switch channel ASC-PRC1_P_SW1R within EZCA_TIMEOUT (2.0s).
2019-09-11_09:24:08.949899Z ALIGN_IFO EZCA CONNECTION ERROR. attempting to reestablish...
2019-09-11_09:24:08.950576Z ALIGN_IFO CERROR: State method raised an EzcaConnectionError exception.
2019-09-11_09:24:08.950576Z ALIGN_IFO CERROR: Current state method will be rerun until the connection error clears.
2019-09-11_09:24:08.950576Z ALIGN_IFO CERROR: If CERROR does not clear, try setting OP:STOP to kill worker, followed by OP:EXEC to resume.
2019-09-11_09:26:16.909062Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP1_P => OFF: INPUT
2019-09-11_09:26:17.160830Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP2_P => OFF: INPUT
2019-09-11_09:26:29.469452Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP1_P => OFF: INPUT
2019-09-11_09:26:29.721182Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP2_P => OFF: INPUT
2019-09-11_09:26:42.038308Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP1_P => OFF: INPUT
2019-09-11_09:26:42.290377Z ALIGN_IFO [DOWN.main] ezca: H1:ASC-INP2_P => OFF: INPUT
....
I then gave up on ALIGN_IFO (I took INIT ALIGN to DOWN a while ago), and moved back to ISC_LOCK and ran a DOWN, but now it looks like it's having the same issue. For isc lock it's listing a connection error...with a different channel & can't perform a down and is stuck in a loop.
Not sure what I can do at this point.......will keep investigating Guardian Land. :-/
Marking this DOWN TIME as CORRECTIVE MAINTENANCE since I can no longer run an alignment...or even return to locking apparently.
Here is the error I get with ISC LOCK when I try to run a DOWN: (and this was after taking OP to STOP, letting it complete, and then taking OP to EXEC)
H1:LSC-REFLBIAS_SW2 => 0
2019-09-11_09:38:48.097459Z ISC_LOCK [DOWN.main] ezca: H1:LSC-REFLBIAS => ON: FM9, FM3
2019-09-11_09:38:50.274110Z ISC_LOCK [DOWN.main] USERMSG 3: EZCA CONNECTION ERROR: Did not observe effect of writing value to switch channel SUS-ETMX_M0_LOCK_P_SW1R within EZCA_TIMEOUT (2.0s).
2019-09-11_09:38:50.277456Z ISC_LOCK EZCA CONNECTION ERROR. attempting to reestablish...
2019-09-11_09:38:50.307871Z ISC_LOCK CERROR: State method raised an EzcaConnectionError exception.
2019-09-11_09:38:50.307871Z ISC_LOCK CERROR: Current state method will be rerun until the connection error clears.
2019-09-11_09:38:50.307871Z ISC_LOCK CERROR: If CERROR does not clear, try setting OP:STOP to kill worker, followed by OP:EXEC to resume.
2019-09-11_09:38:50.618404Z ISC_LOCK [DOWN.main] ezca: H1:ASC-DHARD_P => OFF: INPUT, OFFSET
2019-09-11_09:38:50.870580Z ISC_LOCK [DOWN.main] ezca: H1:ASC-DHARD_Y => OFF: INPUT, OFFSET
2019-09-11_09:38:51.122586Z ISC_LOCK [DOWN.main] ezca: H1:ASC-CHARD_P => OFF: INPUT, OFFSET
2019-09-11_09:38:51.374299Z ISC_LOCK [DOWN.main] ezca: H1:ASC-CHARD_Y => OFF: INPUT, OFFSET
2019-09-11_09:38:51.626122Z ISC_LOCK [DOWN.main] ezca: H1:ASC-DSOFT_P => OFF: INPUT, OFFSET
2019-09-11_09:38:51.888774Z ISC_LOCK [DOWN.main] ezca: H1:ASC-DSOFT_Y => OFF: INPUT, OFFSET
2019-09-11_09:38:52.145125Z ISC_LOCK [DOWN.main] ezca: H1:ASC-CSOFT_P => OFF: INPUT, OFFSET
2019-09-11_09:38:52.396948Z ISC_LOCK [DOWN.main] ezca: H1:ASC-CSOFT_Y => OFF: INPUT, OFFSET
2019-09-11_09:38:52.401265Z ISC_LOCK [DOWN.main] ezca: H1:LSC-REFLBIAS_SW1 => 256
2019-09-11_09:38:52.401855Z ISC_LOCK [DOWN.main] ezca: H1:LSC-REFLBIAS_SW2 => 80
2019-09-11_09:38:52.402348Z ISC_LOCK [DOWN.main] ezca: H1:LSC-REFLBIAS => OFF: FM1, FM2, FM3, FM4, FM5, FM6, FM7, FM8, FM9, FM10
2019-09-11_09:38:52.402773Z ISC_LOCK [DOWN.main] ezca: H1:LSC-REFLBIAS_SW1 => 0
2019-09-11_09:38:52.403250Z ISC_LOCK [DOWN.main] ezca: H1:LSC-REFLBIAS_SW2 => 0
2019-09-11_09:38:52.403250Z ISC_LOCK [DOWN.main] ezca: H1:LSC-REFLBIAS => ON: FM9, FM3
2019-09-11_09:38:54.839541Z ISC_LOCK [DOWN.main] ezca: H1:ASC-DHARD_P => OFF: INPUT, OFFSET
2019-09-11_09:38:55.091397Z ISC_LOCK [DOWN.main] ezca: H1:ASC-DHARD_Y => OFF: INPUT, OFFSET
2019-09-11_09:38:55.343344Z ISC_LOCK [DOWN.main] ezca: H1:ASC-CHARD_P => OFF: INPUT, OFFSET
2019-09-11_09:38:55.595399Z ISC_LOCK [DOWN.main] ezca: H1:ASC-CHARD_Y => OFF: INPUT, OFFSET
Sent out texts of help to Jenne & Jamie.
Jamie luckily was in Europe so it's morning time for him. He's helping me look through the issues currently.
Just for future reference, this was an issue with h1boot1, not guardian.
https://alog.ligo-wa.caltech.edu/aLOG/index.php?callRep=51883
Fil and I went to EX to look inside the BRSX enclosure in prep for the work next month. While we were there I touched up the centering of the BRS. While we had the box open, we must have disturbed the cable into the interface box inside, because when we got back to the corner station the BRS wasn't damping down. I went back down to look at it and it took a little to realize there is not retention on the ethercat cable into the interface box, so it's really easy to knock loose. It immediately started damping down when I pushed the ethernet cable into the interface chassis.
The BRS is still calming down, so we shouldn't use it for a little while yet. The operator should watch the the BRS health ndscope and wait for the BRSX RY OUT to look more like the BRSY RX OUT signal, i.e. the blue trace on the middle timeseries looks like the yellow(?) trace on that same plot.
Opened FRS Ticket 13570 to track that H1 BRS EX ethernet connector at interface box needs to be repaired.
Is it safe to transition from WINDY NO BRSX back to regular WINDY while in NOMINAL LOW NOISE? (Or do we have to wait until we are not locked to make this transition?)
The transition is handled by the SEI_CONF guardian, and is very similar to making the transition to the earthquake state, and does not affect the observation bit. If anything, putting the BRSX back into the loop should make the IFO quieter, assuming the BRS has calmed down.
Looking at minute trends of the sensor correction outputs for EX and EY for the last day, we could have transitioned back about 5pm local last night, first plot.
This is inspite of the fact that the BRSX was still not totally settled, the DC position is still slowly settling down, second plot.
And, the BRSX RY out still had something like a 5000 nrad offset (third image), it was moving slowly enough to not affect the sensor correction signal.
Ah, My Bad.
Could this be the source of my woes with locking after the h1boot1 recovery?