Resolved -
We have brought up all nodes affected by the cooling outage. Users should be able to submit jobs normally.
A few nodes are still down due to unrelated issues, which we will continue to investigate.
Aug 18, 15:01 CDT
Investigating -
Due to an unexpected cooling outage, a majority of the nodes on the HPC cluster are down, and jobs were interrupted. We are working to bring the system back up.
Aug 18, 09:27 CDT
Resolved -
This incident has been resolved.
Aug 11, 08:51 CDT
Monitoring -
A fix has been implemented and we are monitoring the results.
Aug 10, 17:30 CDT
Investigating -
Users are experiencing slow or hanging commands on spark-login. Any users running heavy processes on spark-login should cancel/kill these processes if possible. We are currently investigating.
Aug 10, 16:55 CDT
Resolved -
This incident has been resolved.
Aug 11, 08:50 CDT
Investigating -
Some jobs matching to gpu4002 are evicted immediately with the error message, "stat failed on /var/lib/condor/execute/slot2: No such file or directory". We are currently investigating.
Aug 10, 16:15 CDT