Displaying reports 37181-37200 of 89161.Go to page Start 1856 1857 1858 1859 1860 1861 1862 1863 1864 End
Reports until 08:01, Friday 29 November 2019
H1 General
edmond.merilh@LIGO.ORG - posted 08:01, Friday 29 November 2019 (53570)
Shift Summary - Owl

TITLE: 11/29 Owl Shift: 08:00-16:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Corrective Maintenance
INCOMING OPERATOR: Jim
SHIFT SUMMARY:

Locking Recovery After Computer Re-Boot: Not going so well.


LOG:

EX Saturations

13:54 Superevent S191129u

H1 CDS
david.barker@LIGO.ORG - posted 06:34, Friday 29 November 2019 - last comment - 06:47, Friday 29 November 2019(53564)
h1seib1 crash

Looks like a general core crash, models remain running. Unlike earlier h1seib3 crash, in this case no other systems are affected and only DAQ data from SEI-B1 is down.

Images attached to this report
Comments related to this report
david.barker@LIGO.ORG - 06:41, Friday 29 November 2019 (53566)

I've taken h1seib1 out of the Dolphin fabric, which caused lock loss. Ed is now rebooting the machine via the front panel reset button.

david.barker@LIGO.ORG - 06:43, Friday 29 November 2019 (53567)

Opened FRS13892

david.barker@LIGO.ORG - 06:47, Friday 29 November 2019 (53568)

Computer rebooted with no problems. No IRIG-B timing issue. All ADC and DAC cards seen.

Over to Ed for IFO recovery

H1 CDS (SEI)
edmond.merilh@LIGO.ORG - posted 06:31, Friday 29 November 2019 - last comment - 06:47, Friday 29 November 2019(53563)
H1IOPSEIB1 Crashed

14:12 SEI B1 computer went down and reported ITM trips, but H1 remains locked.

14:22 Contact made with Dave Barker - waiting on a return call

 

Comments related to this report
edmond.merilh@LIGO.ORG - 06:37, Friday 29 November 2019 (53565)

14:36 Dave is removing it from Dolphin to reboot which will likely caue lockloss

edmond.merilh@LIGO.ORG - 06:47, Friday 29 November 2019 (53569)

14:40 Lockloss due to computer work

  • ISC_LOCK set to DOWN

14:42 H1IOPSEIB1 FE computer RESET

14:46 reset SWWD

  • Re-Boot was seemingly routine
  • ready to re-lock

 

H1 General
edmond.merilh@LIGO.ORG - posted 04:27, Friday 29 November 2019 (53562)
Mid-Shift Report

While locked and Observing for 7hours, H1 has been increasingly glitchy with only 2 EX saturations

The wind is up in excess of 20MPH. There are no other obvious causes available on FOMs. Glitching ranges from half spectrum range (lower) to broad range.

H1 General
camilla.compton@LIGO.ORG - posted 00:02, Friday 29 November 2019 (53560)
Shift Summary - Eve
TITLE: 11/29 Eve Shift: 00:00-08:00 UTC (16:00-00:00 PST), all times posted in UTC
STATE of H1: Observing at 115Mpc
INCOMING OPERATOR: Ed
SHIFT SUMMARY: Locked for 2h30
LOG: See previous alogs for relocking notes.
H1 General
edmond.merilh@LIGO.ORG - posted 00:00, Friday 29 November 2019 (53561)
Shift Transition - Owl

TITLE: 11/29 Owl Shift: 08:00-16:00 UTC (00:00-08:00 PST), all times posted in UTC
STATE of H1: Observing at 115Mpc
OUTGOING OPERATOR: Camilla
CURRENT ENVIRONMENT:
    SEI_CONF state: WINDY
    Wind: 10mph Gusts, 8mph 5min avg
    Primary useism: 0.03 μm/s
    Secondary useism: 0.28 μm/s
QUICK SUMMARY:

H1 General
camilla.compton@LIGO.ORG - posted 19:41, Thursday 28 November 2019 - last comment - 21:38, Thursday 28 November 2019(53556)
H1 Lockloss
Comments related to this report
camilla.compton@LIGO.ORG - 20:56, Thursday 28 November 2019 (53557)
  • 03:45 Lockloss #4 at RESONANCE
  • 03:50 Message "IR not found" but then it quickly continued and locked DRMI, maybe I should have been quicker and stopped it to manually adjust IR alignment here?
  • 03:56 Lockloss #5 at SHUTTER_ALS. Just before this lockloss can see ASC_DHARD channels have a strange change in them just before lockloss (photo: lockloss5)
  • 04:08 Lockloss #6 at DARM_TO_RF
    • This time I try to get better IR flashes, adjust COMM and DIFF but still struggle to get consistent DIFF signal.
  • 04:30 Lockloss #7 at CARM_OFFSET_REDUCTION. The ASC_DHARD channels have the same signal as in lockloss 5 (photo: lckloss7). Addionaly, the SUS ETMX UIM/L1 outputs that Jeff adjusted (alog 53540) got up to 140,000 before this lockloss (photo: lockloss7) so if we loose lock aagin I think I will try revetring these changes.
  • 04:34 Started inital alignment
    • INITAL_ALIGN got again stuck at OFFLOADING_GRREN_WFS so I again manually took ALS_X/YARM to UNLOCKED (as worked in alog 53552)
Images attached to this comment
camilla.compton@LIGO.ORG - 21:30, Thursday 28 November 2019 (53558)
  • 05:26 NLN and OBSERVING!!
edmond.merilh@LIGO.ORG - 21:38, Thursday 28 November 2019 (53559)

Good job! And... thank you!

H1 General
camilla.compton@LIGO.ORG - posted 18:14, Thursday 28 November 2019 - last comment - 18:33, Thursday 28 November 2019(53552)
H1 Lockloss
Following on from Jim's alog 53549:
Currently at INCREASE_POWER with my fingers crossed.
Comments related to this report
camilla.compton@LIGO.ORG - 18:33, Thursday 28 November 2019 (53555)
  • 02:25 NLN. Only SDF I accepted was H1:ALS-Y_PZT2_TRIG_THRESH_ON that I changed (photo attached). I don't understand it well enough to know what it effects it has.
  • 02:27 OBSERVING!
Images attached to this comment
H1 CDS
david.barker@LIGO.ORG - posted 08:55, Thursday 28 November 2019 - last comment - 16:57, Thursday 28 November 2019(53547)
comments on symptoms from h1seib3 general core crash

You might have noticed the strange gap in the wind trend ndscope covering the time h1seib3 was down (plot 3). This is because when h1seib3's general cores went offline, the mx_stream process (which sends the front end data to the DAQ) stopped running. This in turn caused problems for all the front ends which share the same ethernet port on the DAQ data concentrator. This is why most of the corner station frontend models went into DAQ data error (plot 2), and why it was critical to fix the problem quickly. One of the models so affected was the external EPICS data collector h1edc running on h1susauxh34 (plot 1). So all slow data was also frozen at the last received value.

 

Images attached to this report
Comments related to this report
david.barker@LIGO.ORG - 09:14, Thursday 28 November 2019 (53548)

TJ mentioned in his alog that the LOCKLOSS_SHUTTER_CHECK.py guardian code had problems reading the HAM6 GS13 channel during the time h1seib3 was down. This was another symptom of the DAQ data being corrupted on most of the corner station models, in this case h1isiham6's data.

This guardian node reads the fast GS13 channel H1:ISI-HAM6_BLND_GS13Z_IN1_DQ from the NDS, which would have been a repeat of the last 62.5mS of data (16Hz). Looking at this channel's raw data would appear to show noise. The attached minute trend plot shows this data repetition over the period h1seib3 was down.

Images attached to this comment
keith.thorne@LIGO.ORG - 16:57, Thursday 28 November 2019 (53551)
Just a note from the future: the new CDS DAQ scheme utilizing a non-Myricom driver for the data concentrator (in full-scale testing from before start of O3), does not suffer this particular pathology. This has even been proven on L1 CDS, where the new scheme is running parasitically to an alternate data concentrator in parallel with the old one.  
H1 CDS
david.barker@LIGO.ORG - posted 19:03, Wednesday 27 November 2019 - last comment - 17:41, Thursday 28 November 2019(53533)
h1seib3 general core crash, models still running but most of Corner Station DAQ data bad

h1seib3 has suffered a general core crash, meaning the models are still running (h1iopseib3, h1isiitmx and h1hpiitmx) but the DAQ data from itself, and most of the other SUS, SEI and PSL front ends are invalid (see attached). The only recourse is to reboot h1seib3. In preparation for this, I took this machine out of the Dolphin fabric (which killed the lock). Jeff is in the MSR taking a photo of the console and rebooting h1seib3 via the front panel reset button.

Comments related to this report
david.barker@LIGO.ORG - 19:04, Wednesday 27 November 2019 (53534)

I've opened FRS13888

jeffrey.kissel@LIGO.ORG - 19:06, Wednesday 27 November 2019 (53535)
Here?s a picture of the rack ?KBM? console before we restarted the computer, as instructed by Dave.
Images attached to this comment
david.barker@LIGO.ORG - 19:19, Wednesday 27 November 2019 (53536)

The computer was powered down by holding the power button, but instead of a clean off-then-on the leds lit up immediately, which suggests the power button may have gotten wedged by a misaligned bezel (and I was not able to ping the machine even after several minutes). Jeff went back into the MSR to free the button, and I quickly re-disabled the IX dolphin port in case the computer had in fact started to reboot, in which case the second reboot would have crashed the entire corner station.

At this point in time h1seib3 has booted and its models are running, but the IOP model has a negative IRIG-B excursion so all IPC and DAQ data are bad. It wont be until the timing is correct will I know if the Dolphin port is still disabled, in which case I will re-enable it.

david.barker@LIGO.ORG - 19:25, Wednesday 27 November 2019 (53537)

h1iopseib3 timing came good, its IPC and DAQ data is now good.

To complete the cleanup, I pressed the DIAG_RESET buttons on all models and cleared h1dc0's CRC counters. The CDS Overview is now green again.

Handing over to Camilla and Jeff for IFO recovery following a SEI-ITMX model reboot.

david.barker@LIGO.ORG - 19:26, Wednesday 27 November 2019 (53538)

Post-Dated Entry: Here is the CDS Overview at the time of the crash:

Images attached to this comment
keith.thorne@LIGO.ORG - 17:41, Thursday 28 November 2019 (53553)
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler.  We have this implemented at LLO - see OpsWiki-Front-ends.

As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here).  The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS.  But we want a lot of full-scale testing done first, so after O3.
keith.thorne@LIGO.ORG - 17:41, Thursday 28 November 2019 (53554)
A power-cycle using the IPMI port (with built-in Dolphin port disable) might have been simpler.  We have this implemented at LLO - see OpsWiki-Front-ends.

As for the OS lock-ups, these likely will not get fixed until we move to a newer Linux kernel (3.0.8 here).  The small test systems that ran it for several years just did not provide the statistics to reveal these problems. The good news is we had a breakthrough this summer and have real-time systems running on current kernels (4.19 - Debian 10) patched for CDS.  But we want a lot of full-scale testing done first, so after O3.
Displaying reports 37181-37200 of 89161.Go to page Start 1856 1857 1858 1859 1860 1861 1862 1863 1864 End