On Tuesday 23rd April I turned on a ZFS scrub for the /opt/rtcds file system and we got a skew of h1edc CRC errors until I stopped the scrub.
To investigate this further, I wrote a script which creates a 1GB file on both the /ligo and /opt/rtcds file systems from a workstation. This runs every 5 minutes and writes the disk-access speed into two EPICS records (units = MB/s).
At 04:49 PDT this morning we got 4 CRC errors, which in this case correlated to a slow down of /opt/rtcds access speed of 30%.
Attached plot shows past 48 hour trend of h1edc CRC (upper ndscope plot) and file access speed (lower strip-tool) which have been aligned on their x-axis. The strip tool shows /opt/rtcds access speeds (blue line) and /ligo (green line).
Note:
Both /ligo and /opt/rtcds slowed down this morning between 2am and 9am, with /opt/rtcds significantly slowing around 5am.
The previous CRC error of 3 counts from almost 48 hours ago is not related to disk slow down.
/opt/rtcds/ also slowed between 4am and 6am Thursday morning.
Jonathan checked that network statistics between the hours of 2am and 9am this morning, nothing unusual was found (core switch showed a slight traffic increase between 2am and 4am).
Jonathan, Dave:
One way a NFS file system slow down could impact h1edc is during its periodic calculation of the file-checksum for the file H1EDC.ini. To further test this, Jonathan made a RCG-branch3.5 change on the DTS to take this checking out of x1edc (alog link below). This has been running for 90mins with no errors. Overnight last night x1edc raised 700+ CRC errors, so by Monday we should have a definitive result.
https://alog.ligo-la.caltech.edu/TST/index.php?callRep=12509