Search criteria
Section: H1
Task: CDS
WP13564 SPI Added QPD Filter modules and DQ chans
Jeff, Jennie, Jonathan, Dave:
We installed the new h1spih23 model at 12:15. The DAQ was restarted soon afterwards. 0-leg 12:16 and 1-leg 12:20.
ndscope is updated to 0.22.0 on all the workstations. Highlights include:
* Fixed bug that prevented showing some fast channels at full resolution
* Use shift-drag to constrain move/zoom to the time axis and use ctrl-drag to constrain move/zoom to the Y axis
* Allow full precision when entering scale & offset for a channel
* New plots now get custom trace width when set
* Use '--replay-id ' to access Arrakis replays
* Button to add Y cursors to all plots
Full change log here: https://git.ligo.org/cds/software/ndscope/-/blob/master/CHANGELOG.md#0220---2026-08-24
Jonathan found a correlation between the number of NDS requests around 15:15 this afternoon and the crash of nds1-daqd. Plot show number of processes and cpu usage against nds1-daqd uptime. Logs suggest a nds client rapidly asked for data, possibly on cdsws37.
WP 13532. Dave, Gerardo, Jordan, Patrick The EtherCAT configuration and PLC code on h0vaclx have been updated to change the PT180 gauge on BSC8 from a BCG 450 to a BCG 552. In the process I found out that PT140 had been removed in the past month or so, so I also removed it from the configuration and code. I also had errors from PT193, and learned that it was potentially broken, and along with PT191 and PT192 was slated for removal. I therefore asked Jordan to disconnect these from the EtherCAT hub, and removed them from the EtherCAT configuration and PLC code as well. When I tried to take the system to OP, the second section of the CU1128 EtherCAT hub kept going into an error state. I eventually tracked it down to what I think was a mismatch between the vendor id of the hub in the TwinCAT solution and that of the actual hardware. Oddly I don't remember this being an issue in the past. I could find no indication of what the vendor id was in the solution, other than what the error was reporting, or how to change my script to make it match. I found a way around this by selecting and choosing 'change to compatible type' for each of the three parts of the hub in the generated solution, rescanning the IO tree, and when it showed mismatches in the version numbering, used the dialog box to change what was configured to what was found by the scan. This seemed to fix the problem. Dave restored the PID and other settings from SDF. The channels have been updated in the DAQ. The tripped high voltages have been turned back on. I have closed the work permit.
WP13536 SUS BS additional fast channels
Oli, Dave:
We installed Oli's latest h1susbs which added 13 new fast DQ channels, most at 2k, some at 256Hz. h1susbs model was restarted at 12:37 and the DAQ soon after.
WP13535 Add DAQSTAT channels to DAQ
Dave:
The new DAQSTAT's non-string records were added to the DAQ as EDC channels.
WP13542 Remove h1daqscript0 epics-load-mon channels from DAQ
Dave:
A new H1EPICS_CDSMON.ini was generated sans DAQSCRIPT0 channels. DAQ restart was needed
WP13532 Upgrade VAC LX Beckhoff and IOC, update PT180, remove obsolete gauges.
Patrick, Dave:
Patrick installed the latest VAC LX code and I installed its new INI file into the DAQ.
Gauges removed: PT140, PT191, PT192, PT193.
PT180 was modified, some channels removed, some added, some stayed the same name.
Note that PT180's MOD1 and MOD2 gauges have been reversed
| Before Today | Now | |
| MOD1 | Had been BSC8 gauge, recently has been reading zero | reads ~900 |
| MOD2 | Had been reading ~740 | Now is BSC8 gauge, reading 7e-08 Torr |
DAQ Restart
Jonathan, Dave
We did a DAQ restart, 1-leg and EDC at 12:39, 0-leg at 12:41. Frame writers (which are not restarted with the new code) had caught up with the new run number and were writing identical full frames at 12:44.
At 15:16 NDS1 restarted itself. Jonathan is investigating.
Per WP 13538 we updated and restarte the IOC container cluster. This is getting more visibility as we have moved more of our soft IOCs to the cluster and the brief outages may be noticed when we update and reboot. So for the next few cycles at least this will be done under work permit. Part of the goal of this infrastructure work is to be possible to update these hosts (and to be more discoverable, and resilent).
Workstations were updated and rebooted. This was an OS packages update. Conda packages were not updated.
EJ, Dave:
We got another "2nd ADC timing error" event last night on h1iopoaf0. The previous one was Wed 11aug2026 on h1iopiscex (alog 91508). Unlike the h1iscex one, h1iopoaf0's cpu did not spike, but it is running high anyway. No dmesg log was generated.
It is looking more likely this is software and not an issue with the second Adnaco backplane as postulated in the 11th Aug alog.
Oli, Erik, Dave
Oli found that tconvert was complaining of an expired /ligo/data/tcleaps/tcleaps.txt file. The "valid through" line in this file is the UNIX time for 07aug2026, which is when we think this warning started.
We have extended the "valid through" time by 10 years, it will expire August 2036.
SQZ slow controls SDF became unresponsive late this morning. I tried a restart of this model on h1ecatmon0
> rtcds restart h1syscssqzsdf
but it did not read the monitor request file and ran with zeros in all the channel counts.
After verifying the monitor.req and safe.snap files existed and looked ok, I did a slow stop/start of this model with about 1 minute inbetween. This time it read the monitor file and is monitoring 1340 chans.
I've completed the upgrade of DAQ stat. New features ported from the old system are:
- DC broadcast list check-sums must be identical
- DC FE network port bandwidth must match within a tolerance
- FWs disk usage is displayed, warnings issued to too large or too small
The MEDM was getting too tall to fit a 2k monitor, so the framewriter section was moved to the right.
Remaining task is to add the new channels to the DAQ next week.
Claude has written a DAQSTAT User Guide, which I've posted to the DCC as T2000607
Daniel tried to rebooted Beckhoff PC remotely, but it didn't come back so he rebooted it in MSR.
After that things came back mostly, there were two major exceptions:
1. JAC WFS whitening/dewhitening mismatch.
WFS whitening was off for two stages we use but the anti-whitening was on. WFS was at first turned on by the guardian when we tried to lock JAC and it totally misaligned it. We had to manually locked JAC, noticed that the alignment was off, cleared WFS history both in WFS feedback filters as well as JM1 lock filters, fixed the whitening/dewhitening mismatch by pressing "ON" buttons for two whitening stages, and after that it started working OK.
2. JAC got cooler, temperature servo overheats at first, and we're only getting 15 minutes lock stretches for now.
It took about 20 minutes for the Beckhoff to come back because the first reboot attempt didn't work (1st attachment). Thermistor 1 temperature dropped by about 0.05 degC as the temperature servo wasn't doing anything (2nd attachment, it looks like a sudden change but it's not, it's just that the Beckhoff EPICS numbers were frozen before that). Temperature servo started heating it hard as soon as it came back, overshot and tried to reduce heating, but it's still heating harder than it used to. Because of that, JAC lock PZT voltage quickly goes to zero after relocking. We'll have to wait.
Jonathan, Dave:
daqstat was the last service to move off of h1daqscript0 (others were Dolphin network manager and daq_run_number_server). We powered this machine down at 16:30 today. I have temporarily greened the EDC by adding h1daqscript0's epics_load_mon channels to edc_green_ioc.
The LHO DAQ has three frame-writers. Two (FW0 and FW1) write to LDAS and have large disk arrays, the third (FW2) runs with a local disk and is used as a "tie breaker" if the two primary machines disagree with each other.
There are now 5 permutations of agreement between the 3 frame-writers:
All agree (0=1=2)
All disagree (0!=1!=2)
One disagrees with the other two ((0=1)!=2) ((0=2)!=1) and ((1=2)!=0)
Instead of a simple red/green LED indicator on the MEDM, we now have borders around each monitor which are red/green depending upon the agreement.
This post made at 09:50 PDT.
Jonathan is working on the issue.
Miranda, Dave:
Over the past week we have documented the LHO PEM accelerometers and produced an as-built wiring drawing
This was in response to a possible misreading of accelerometer(s) with the wrong channel names.
We found three accelerometers with incorrect names, which date back to pre-O4, so I'm tagging DETCHAR.
| The accelerometer | Was being readout as this channel |
| H1:PEM-CS_ACC_LVEAFLOOR_HAM6_Z_DQ | H1:PEM-CS_ACC_BEAMTUBE_SRTUBE_X_DQ |
| H1:PEM-CS_ACC_BEAMTUBE_SRTUBE_X_DQ | H1:PEM-CS_ACC_EBAY_FLOOR_Z_DQ |
| H1:PEM-CS_ACC_EBAY_FLOOR_Z_DQ | H1:PEM-CS_ACC_LVEAFLOOR_HAM6_Z_DQ |
We fixed this at 16:00 Thursday 13th August 2026 PDT by remapping the BNC connections on the front panel of APC1. The cable going to port12 was moved to port13, that going to port13 was moved to port14 and that going to port14 was moved to port12.
We tap-tested the EBAY_FLOOR accelerometer at 16:10 and this showed up on H1:PEM-CS_ACC_EBAY_FLOOR_Z_DQ
For detchar: Based on our (dave, robert, my) best guesses, this error starts on JAN 26th 2024 when the ENdevcos, as well as the signal conditioners for the accelerometers, were replaced (alog 75592). This is post end of O4a (Jan 16th 2024), so the data for O4a is okay, but the data for O4b/c is affected. This means that if there's an hveto round winner/trigger/pemcheck correlation/etc for one of the three channels on the right, it will instead be data from the detector on the left column. For clarity:
Here are some examples that I hope help clarify (as I was also confused):
Dave noticed that the switch on the power conditioner for H1:PEM-CS_ACC_HAM7_Y was set to x1 (visible in the wiring diagram, v6), while all the other switches were set to x10. Today, Robert and I went in and switched it to x10 at 11:30 AM PDT. We both believe this was related to the ENdevco replacements on Jan 26th, 2024. I looked at the spectra before and after replacement (Jan 24th, 2024 and Jan 29th, 2024, attached below), and then today after we flipped it. As you can see, the spectra now has the correct order of magnitude.