EJ, Erik, Jonathan, Tony, Dan M,
This morning EJ and Jonathan started models on the systems and did a daqd restart as rcg5.6 adds a few new channels. We found that there were install problems with h1suslo12 and h1susauxh6. EJ looked into those and fixed some permission issues and was able to re-install the models.
We had problems with lsc, seib2, seih23
* LSC - EJ noted that the LSC system had a correctable pci error yesterday around 5:25pm localtime. We were not controlling anything at this point. The LSC models would not stay running. EJ fixed some BIOS setting errors (the c-state settings were not locked to c0/c1 state) and enhanced fault state was not disabled. EJ traced this down to spikes in the model run time to over 400ms! We when through and tried a real time processing benchmark with no add in IO cards, the new ethernet NIC, the adnico cards (w/o fibers plugged in), the adnico cards with fibers, and then variations on which fibers. In the end we have left the 4th adnaco disconnected. We suspect the fiber got dirty as it was working. There are no io cards in that chassis on that card.
* seib2 did not see any of its binary io cards. Rebooting was sufficient to fix that.
* seih23 did not see any IO cards, a reboot fixed that.
After fixing the lsc model Erik and Tony took a spare computer out to end-X to look at seiex. They were unable to get the computer to see the new ethernet card. They ended up putting the spare computer in. Once that was done seiex worked.
After seiex was running, EJ looked at the global IPC state. We had to restart the ethernet controller and models on cdsh8, as it was erroring out on all its ipcs. The h1omcpi was having a steady IPC error rate of about 20 errors a second. The errors are coming from the h1ioplsc0 model. This is a problem with the fast adc logic just taking a lot of time. The IPCs where coming in late for the omcpi model. EJ was able to adjust how it waited for the IPCs to clear the errors. The base assumption on the IPCS timings is against models that are running with some headroom, the lsc iop mode is running at almost the max time that it can which gives a small window for the IPC.
We restarted h1build and h1ecatmon0 to have them booting from the new bootserver.
Dan Moraru and Jonathan looked at the cdsfs systems. Dan was able to fix a race condition which prevented the system from doing high availability fail over for the /ligo filesystem. Now we can move which server is serving /ligo between cfdsf2 and cdsfs3. Migrating the /ligo filesystem causes access to the filesystem by the control room systems to pause for a few minutes. We tested this transition a few times, and then tested it by applying system updates and reboots on cdsfs2 and cdsfs3.
Dan also put the same configuration fixes into cdsfs4 and cdsfs5. These are designed to hold /opt/rtcds (though they are not yet). We will continue to work on these systems as the cdsfs4 & cdsfs5 have not come back properly from updates.