智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 09-07

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article算电协同

Hosting Capacity Assessment of Data Centers with Voltage Ride-Through Capability in Power Systems

Pengyu Ren、Wei Sun、Fei Teng

Published 2026-09-03 · arXiv · Credibility S

Large data centers are emerging as concentrated, power-electronic grid loads whose abrupt disconnection or transfer to on-site backup supply during voltage disturbances can remove large demand from the power system, and may create a system-level stability problem. Their interconnection feasibility therefore depends not only on steady-state thermal and voltage limits, but also on whether internal power-conditioning s…

Abstract, interpretation and reference

Abstract

Large data centers are emerging as concentrated, power-electronic grid loads whose abrupt disconnection or transfer to on-site backup supply during voltage disturbances can remove large demand from the power system, and may create a system-level stability problem. Their interconnection feasibility therefore depends not only on steady-state thermal and voltage limits, but also on whether internal power-conditioning systems can maintain IT service while limiting customer-initiated load reduction. This paper presents a voltage ride-through (VRT)-aware data center and grid co-planning framework that couples transmission-level fault simulation with an internal data center ride-through model. Python-based dynamic simulations generate point-of-interconnection (POI) voltage trajectories under selected network faults, and the resulting waveforms drive an internal model incorporating IT and cooling-load dynamics, DC-link, Uninterruptible Power Supply (UPS) response, and converter apparent power limits. The IEEE 118-bus case study shows that internal VRT capability can become a binding interconnection constraint: steady-state planning alone can overestimate feasible data center capacity, whereas increased UPS converter headroom progressively restores hosting capacity. Under the reduced-order response models studied, the grid-forming mode provides greater ride-through margin than the current-limited grid-following mode under the same network fault conditions. The results further show that VRT constraints can materially change both the total hosting capacity of data centers and its spatial allocation across candidate interconnection buses.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Pengyu Ren, Wei Sun, Fei Teng. Hosting Capacity Assessment of Data Centers with Voltage Ride-Through Capability in Power Systems[J/OL]. (2026-09-03)[2026-09-07]. http://arxiv.org/abs/2609.03030v1.

Full text 中文海报
算电协同 论文图示
Research Article芯片与算力

ClusterBench: A Framework for Cluster-Wide Continuous Benchmarking and Regression Testing

Aditya Ujeniya、Jan Eitzinger、Thomas Gruber、Georg Hager、Gerhard Wellein

Published 2026-08-11 · arXiv · Credibility S

Data centers need tooling that validates an entire installation rather than individual nodes, at acceptance and at regular intervals thereafter. This requires dispatching identical benchmarks to every node in a single submission, and therefore cluster-aware scheduling. This paper presents ClusterBench, a framework for cluster-wide continuous benchmarking. It ships with a benchmark collection targeting each component…

Abstract, interpretation and reference

Abstract

Data centers need tooling that validates an entire installation rather than individual nodes, at acceptance and at regular intervals thereafter. This requires dispatching identical benchmarks to every node in a single submission, and therefore cluster-aware scheduling. This paper presents ClusterBench, a framework for cluster-wide continuous benchmarking. It ships with a benchmark collection targeting each component: CPU, GPU, memory, interconnect, and I/O. Because measurements are repeated throughout the cluster's lifetime, ClusterBench collects data across space and time. Comparison against earlier runs detects performance regressions introduced by software changes, such as kernel updates or new library versions. The measurements also form a dataset for research on hardware variability. On the NHR@FAU clusters Helma, Alex, and Fritz, variation within a single component stays within 1%. Variation across specimens reaches 5%, despite nodes identical by specification. Correlating performance with power draw, frequency, and temperature shows that this relationship differs between air- and liquid-cooled nodes.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,芯片、服务器和高密度算力部署正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向跨地域数据中心负载与电力资源之间的调度关系。意义:对日报读者而言,它可用于判断芯片路线和服务器密度变化如何传导到机房设计。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Aditya Ujeniya, Jan Eitzinger, Thomas Gruber, 等. ClusterBench: A Framework for Cluster-Wide Continuous Benchmarking and Regression Testing[J/OL]. (2026-08-11)[2026-09-07]. http://arxiv.org/abs/2608.10956v1.

Full text 中文海报
芯片与算力 论文图示
Research Article算电协同

Exploiting the Benefits of V2B Application on Peak Shaving of Data Center Loads

Arya Joshi、Hamed Haggi、Chinmay Morankar

Published 2026-09-01 · arXiv · Credibility S

The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cool…

Abstract, interpretation and reference

Abstract

The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cooling loads. Initially, a literature review was conducted considering V2B constraints and optimization methods including SoC limitations, EV participation, tariffs, and building loads. This analysis was then used to develop a conceptual case study of a 10 MW data center in Loudoun County, VA by simulating a temperature-dependent load profile and adjusting the V2B participation of 40 commercial and passenger EVs. Simulation results indicate that, depending on seasonal variations in cooling load demands, strategic deployment of V2B assets between 12-5pm can offset gross cooling loads by 13-36%.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用综述归纳和指标比较,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Arya Joshi, Hamed Haggi, Chinmay Morankar. Exploiting the Benefits of V2B Application on Peak Shaving of Data Center Loads[J/OL]. (2026-09-01)[2026-09-07]. http://arxiv.org/abs/2609.00204v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Flexible Training Workloads in Large-Scale AI Data Centers for Transient-Stability Support in Transmission-Constrained Power Systems

Jae-Kyeong Kim

Published 2026-08-31 · arXiv · Credibility S

The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support.…

Abstract, interpretation and reference

Abstract

The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support. To this end, this paper proposes training-induced load surge (TILS), a fast demand-side strategy that initiates or resumes flexible AI training workloads after fault clearing to increase active-power demand at electrically effective locations. The resulting load increase allows accelerating generators to supply additional electrical power, thereby reducing the accelerating-power imbalance and limiting the first-swing rotor-angle excursion. The underlying mechanism is first clarified in a single-machine infinite-bus (SMIB) system and then evaluated in the IEEE 39-bus system and a large-scale Korean power system. Results across all three systems demonstrate that TILS can increase the transient-stability-constrained generation limit. Larger responses, earlier activation, and siting at buses with a stronger electrical influence on the critical generators provide greater generation-limit increases. These results suggest that the upward load-response capability of AI data centers can provide complementary transient-stability support when sufficient electrical headroom, flexible workloads, and reliable grid-triggered activation are available.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用文献摘要中的模型、实验或案例分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Jae-Kyeong Kim. Flexible Training Workloads in Large-Scale AI Data Centers for Transient-Stability Support in Transmission-Constrained Power Systems[J/OL]. (2026-08-31)[2026-09-07]. http://arxiv.org/abs/2608.30901v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Environmental and Economic Implications of Artificial Intelligence Data Centers in the United States

Johanna Bolaños-Zuñiga、Alberto J. Lamadrid

Published 2026-08-11 · arXiv · Credibility S

In this study, we use electricity demand growth, cooling requirements, and backup system operation to evaluate the environmental and economic implications of artificial intelligence data centers in the United States. Our results indicate that impacts are not determined solely by facility design, but by the broader electricity, water, and land-use systems in which these facilities operate. Emissions are primarily dri…

Abstract, interpretation and reference

Abstract

In this study, we use electricity demand growth, cooling requirements, and backup system operation to evaluate the environmental and economic implications of artificial intelligence data centers in the United States. Our results indicate that impacts are not determined solely by facility design, but by the broader electricity, water, and land-use systems in which these facilities operate. Emissions are primarily driven by electricity consumption and therefore depend on marginal generation mixes, transmission constraints, and the spatial and temporal distribution of demand. Analysis further shows that local effects include pressures on water resources, increased noise exposure, and land-use changes, with outcomes varying across regions and infrastructure conditions. The assessment of technological and operational measures shows that improvements in energy efficiency, cooling configurations, and operational strategies can reduce these impacts, although their effectiveness depends on system-level conditions. Evaluation of regulatory and market structures suggests that existing frameworks may not fully account for location- and time-specific externalities. These findings support the need for integrated policy approaches that align data center deployment and operation with electricity system characteristics, water availability, and land-use planning to improve overall environmental and economic performance.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Johanna Bolaños-Zuñiga, Alberto J. Lamadrid. Environmental and Economic Implications of Artificial Intelligence Data Centers in the United States[J/OL]. (2026-08-11)[2026-09-07]. http://arxiv.org/abs/2608.09882v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

Nicoletta Tsiopani、Moysis Symeonides、George Pallis、Marios D. Dikaiakos

Published 2026-08-13 · arXiv · Credibility S

The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We present InFactPlanner, …

Abstract, interpretation and reference

Abstract

The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We present InFactPlanner, a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites. InFactPlanner combines query traces, hardware-model profiles, candidate site configurations, PUE/WUE parameters, renewable generation models, and time-varying grid carbon intensity to estimate power, energy, carbon emissions, water use, latency, and server utilization. The framework abstracts low-level serving effects into configurable hardware-model profiles, enabling rapid comparison of site selection, capacity placement, hardware, model, renewable integration, and routing choices. We validate the energy accounting pipeline by reproducing reference LLM inference energy estimates with less than 10% deviation, evaluate scalability across multiple data centers and server counts, and demonstrate scenario-driven decision analyses for hardware selection, renewable placement, geographic deployment, and carbon-aware routing. Our results show that sustainability-optimal choices can differ from latency-optimal ones, and that the carbon value of deployment depends strongly on the local grid mix.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Nicoletta Tsiopani, Moysis Symeonides, George Pallis, 等. InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers[J/OL]. (2026-08-13)[2026-09-07]. http://arxiv.org/abs/2608.12915v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Techno-Economic Boundary Analysis of Small Modular Reactor Cogeneration for Hyperscale Data Center IT and Cooling Loads

Honglin Li、Buxin She、Jie Zhang

Published 2026-08-11 · arXiv · Credibility S

Hyperscale data centers are adding firm, high-utilization demand faster than grids can serve it, renewing interest in colocating them with small modular reactors. Such a plant could earn revenue in two ways, selling low-carbon power and diverting steam to absorption chillers that serve a cooling load accounting for 20-40% of facility electricity use, but neither revenue stream has been priced across the conditions t…

Abstract, interpretation and reference

Abstract

Hyperscale data centers are adding firm, high-utilization demand faster than grids can serve it, renewing interest in colocating them with small modular reactors. Such a plant could earn revenue in two ways, selling low-carbon power and diverting steam to absorption chillers that serve a cooling load accounting for 20-40% of facility electricity use, but neither revenue stream has been priced across the conditions that must coincide. Here we co-optimize reactor dispatch, steam extraction, absorption cooling and grid exchange hourly for a 200 MW$_\mathrm{e}$ data center in the Electric Reliability Council of Texas (ERCOT) region, across 109 runs spanning capital, market, policy, financing and cooling efficiency. At 2023 mid-range reactor capital, the nuclear configurations cost 49-62% more than grid supply even with the Section 45Y production tax credit. The viable region opens near \$5,000 kW$_\mathrm{e}^{-1}$, and nth-of-a-kind capital makes them 77-89% cheaper in 2023, though between parity and 34% more expensive in the low-price 2024 market. A carbon price of \$53-64 tCO$_2^{-1}$ closes the mid-range gap under hourly export crediting. Absorption cooling is dispatched in response to hourly electricity prices and supplies 38% of annual cooling, at an added cost of \$9.2 million yr$^{-1}$ relative to the reactor-only plant; that gap closes at an installed absorption cost of \$60 kW$_\mathrm{c}^{-1}$ at baseline efficiency and \$570 kW$_\mathrm{c}^{-1}$ on a legacy-efficiency campus, against surveyed commercial prices of \$450-1,200 kW$_\mathrm{c}^{-1}$. Together these results delineate the capital, market and policy conditions under which colocated reactor cogeneration is competitive with grid procurement, and the range over which each condition moves the outcome.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用综述归纳和指标比较,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Honglin Li, Buxin She, Jie Zhang. Techno-Economic Boundary Analysis of Small Modular Reactor Cogeneration for Hyperscale Data Center IT and Cooling Loads[J/OL]. (2026-08-11)[2026-09-07]. http://arxiv.org/abs/2608.10999v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

AIDC Microgrid Vulnerability Assessment Under Computing-Power Coordinated Attacks

Ze Yu、Hongwei Zhen、Chao Shen、Mingyang Sun

Published 2026-08-11 · arXiv · Credibility S

The rapid growth of large language model (LLM) services is expanding AI data centers (AIDCs), increasing electricity demand and associated carbon emissions. Renewable energy integration can mitigate these impacts but also strengthens the coupling between AIDC loads and inverter-interfaced generation, creating cross-domain cyber-physical vulnerabilities. Specifically, adversarial AI requests alter AIDC power demand, …

Abstract, interpretation and reference

Abstract

The rapid growth of large language model (LLM) services is expanding AI data centers (AIDCs), increasing electricity demand and associated carbon emissions. Renewable energy integration can mitigate these impacts but also strengthens the coupling between AIDC loads and inverter-interfaced generation, creating cross-domain cyber-physical vulnerabilities. Specifically, adversarial AI requests alter AIDC power demand, whereas inverter control tampering modifies source-side dynamics, and their combined impact on system stability varies with generation forecast and demand response uncertainties. To this end, we propose an uncertainty-aware AIDC microgrid vulnerability assessment framework under computing-power coordinated attacks. First, the framework maps adversarial AI requests to AIDC power variations and represents uncertainties in attack-induced demand responses and photovoltaic (PV) forecasts through confidence-weighted realizations. Then, impedance based stability analysis combines these realizations with bounded inverter parameter tampering to construct attack reachable domains and identify critical attack time windows. Furthermore, a separate criterion identifies fixed coordinated attack vectors that retain destabilizing capability throughout each selected window. Case studies demonstrate that, unlike either attack component applied alone, coordinated attacks within identified critical windows induce sustained inverter frequency oscillations with peak absolute deviations exceeding 20% of nominal frequency, whereas the evaluated out-of-window response remains bounded. The proposed method further identifies critical attack windows and the associated coordinated attack vectors.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Ze Yu, Hongwei Zhen, Chao Shen, 等. AIDC Microgrid Vulnerability Assessment Under Computing-Power Coordinated Attacks[J/OL]. (2026-08-11)[2026-09-07]. http://arxiv.org/abs/2608.10645v2.

Full text 中文海报
算电协同 论文图示