Wang, Bo; Müller, Matthias S. (Thesis advisor); Ludwig, Thomas (Thesis advisor); Lakemeyer, Gerhard (Thesis advisor)
Aachen : RWTH Aachen University (2026)
Dissertation / PhD Thesis
Dissertation, RWTH Aachen University, 2026
Abstract
The demand for computational power in both traditional scientific computing and artificial intelligence has continued to increase. Consequently, high-performance computing (HPC) clusters have grown in both size and power dissipation over the past several decades, as evidenced by historical TOP500 records. Over the past twenty years, the computational performance of the flagship TOP500 system has increased by a factor of 60, while its power consumption has nearly doubled. Current and future large-scale HPC clusters face a power wall imposed by the limited capacity of supporting infrastructure and rising energy costs. Therefore, effectively managing and optimizing cluster power consumption has become increasingly important to alleviate infrastructure constraints and reduce energy consumption. In this thesis, first, I discuss a cluster’s total costs of ownership (TCO) by covering one-time and annual costs. The one-time costs arise in a cluster construction stage through investments into infrastructure and compute nodes. In contrast, the annual costs arise from the electricity bill, et cetera. Based on TCO data collected from two real-world clusters in Germany, the costs related to power have a considerable TCO share. Next, I introduce advanced power management to reduce the TCO. However, such management comes with side effects on the clusters’ computational performance due to the throttling of underlying hardware. Therefore, I define cluster productivity to investigate the trade-offs between power and performance. At the cluster construction stage, I define productivity as the amount of installed compute nodes to the total investment. I optimize productivity through improved financial budgeting into the cluster and infrastructures. In this case, I introduce strategies to manage the cluster’s actual power dissipation to remain under a certain power limitation. During the cluster’s operation, I define productivity as the amount of executed user jobs to the TCO. I showcase approaches for reducing energy consumption and costs. Additionally, I present techniques that maximize the job throughput. In order to efficiently manage the power consumption of a large-scale cluster in the above two cases, complex and portable approaches are required to fulfil the requirements of a large number of nodes and diverse running jobs. To tackle these challenges, I introduce a management hierarchy consisting of cluster-wide, job-level, and node-level agents. The cluster agent controls the power management of the entire cluster. It distributes power budgets to running jobs and compute nodes. The job agent regulates the power setting of nodes employed by a job. The node agent manages the power dissipation of distinct hardware components. For each agent, I explore approaches to reduce energy consumption and maintain a power cap, respectively. Finally, I evaluate diverse power-management approaches using productivity models and real-world data. The management is validated to be effective where the TCO is reduced and the productivity is improved.
Institutions
- Chair of High Performance Computing (Computer Science 12) [123010]