EUPEX parter CINI (University of Bologna) had a second poster accepted at the 19th ACM International Conference on Computing Frontiers, which was held from 17 to 19 May 2022 in Turin (Italy).
Abstract: Modern scientific discoveries are driven by an unsatisfiable demand for computational resources. To solve large problems in science, engineering, and business, data centers provide High-Performance Computing (HPC) systems with aggregation of the computing capacity of thousand of computing nodes. Anomaly prediction is critical in order to preserve the continuity of the service of HPC systems and prevent hardware deterioration. In the datacenter, a thermal anomaly occurs when the balance of cooling capacity and computational demand is disturbed. Moreover, this is identifiable from a suspicious/abnormal pattern in the monitoring signals.
In this poster, the anomaly prediction task in the HPC systems is investigated by defining complex statistical rules-based and Deep Learning DL-based anomaly detection methods, then utilizing these anomaly detection methods in an anomaly prediction framework.
Authors:Mohsen S. Ardebili, Andrea Bartolini, Luca Benini
DOI:10.1145/3528416.3530864