High GPU utilization with no processes running
We noticed a while back that several of our GPU cards retained high utilization even though no processes were running on them. nvidia-smi Fri Mar 23 13:52:55 2018 +-----------------------------------------------------------------------------+ | NVIDIA-SMI 375.26 Driver Version: 375.26 | |-------------------------------+----------------------+----------------------+ | GPU Name…
Update feedback
The beegfs cluster was updated to 2015.03.r23. The fhgfs volume is back up and mounted on all nodes. During the Infiniband switch firmware upgrade an error was encountered. We have logged a support call with Mellanox regarding this. The switch …
Maintenance slot 13 March
The BeeGFS (fhgfs) cluster will be offline on Monday 13th March from 09:00 to 17:00 for a major update. Please ensure that all jobs referencing the BeeGFS volume, /researchdata/fhgfs, are completed before 09:00. The firmware on the Mellanox Infiniband switch …
Mellanox MLNX_OFED_LINUX driver update issue
This is an updated entry for the issue we encountered last year upgrading our HPC servers and Infiniband drivers. An updated installation ISO needs to be created that allows kernel support for the newly updated kernel. To create the ISO…
New GPU cards
We have completed the upgrade to the GPU portion of the hex cluster: – Installed new GPU004 server with two nVidia K40 cards. – Two additional nVidia K40 cards added to GPU003. This brings the number of GPU cards in…
High memory nodes downtime
Two of the high memory nodes, 801 and 802, are being moved to a new rack. They will be unavailable until 14:00 7th July.…
June maintenance slot
ICTS will be conducting power maintenance in their data centers on Sunday the 26th of June between 09:00 and 17:00. The Bremner data center will be shut down completely and hence the Hal Slurm cluster will be offline. We will…
HPC January maintenance
The ICTS hex cluster will be down for scheduled maintenance from Monday January 11th 09:00 to Tuesday January 12th 17:00. The head node, data node and all worker nodes will be patched and rebooted, hence all jobs should be canceled…
