智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 07-29

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article热管理与液冷

Assessing Risks of Hydro-Generator Shaft Fatigue from Data Center Load Oscillations

Kaustav Chatterjee, Meghana Ramesh, Shuchismita Biswas, Brett A. Ross, Antos C. Varghese, Sameer Nekkalapu, Slaven Kincic

Published 2026-07-15 · arXiv · Credibility S

Large AI data center loads can introduce persistent sub-synchronous active-power oscillations that may impact nearby generators by exciting torsional modes and increasing shaft stress. This paper presents a model-based framework for evaluating hydro-generator shaft fatigue risk under oscillatory loading. An electromagnetic transient simulation model is developed using a two-mass turbine-generator shaft representatio…

Abstract, interpretation and reference

Abstract

Large AI data center loads can introduce persistent sub-synchronous active-power oscillations that may impact nearby generators by exciting torsional modes and increasing shaft stress. This paper presents a model-based framework for evaluating hydro-generator shaft fatigue risk under oscillatory loading. An electromagnetic transient simulation model is developed using a two-mass turbine-generator shaft representation with parameters from real-world generation units and a configurable AI data center load. The risk assessment is performed in two stages. First, a network transfer function quantifies the propagation of load oscillations from the data center point of interconnection to the hydro-generator terminal. A plant transfer function then characterizes the resulting shaft torque amplification. A frequency-scan approach identifies resonance regions and evaluates torque amplification at individual forcing frequencies. Parametric studies show that amplification is strongly affected by generator-to-turbine inertia ratio and torsional damping. Lower inertia ratios shift torsional modes to lower frequencies and increase amplification, indicating that some Kaplan-type units may be more susceptible than comparable Francis or Pelton units. Reduced damping further increases resonant response and fatigue exposure. A simplified fatigue assessment based on S--N curves and the Goodman diagram relates simulated torque response to mechanical integrity. The resulting Goodman safety factor provides a practical metric for evaluating the impact of persistent AI data center oscillations on hydro-generator service life and supports interconnection studies, oscillation limits, and plant-level monitoring strategies.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures

Soham Ghosh, Nabil Mohammed, Mohammad Ashraf Hossain Sadi

Published 2026-07-19 · arXiv · Credibility S

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operati…

Abstract, interpretation and reference

Abstract

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operational impacts of AI data centers, as well as their potential role as grid-interactive assets, limited attention has been given to the challenges associated with their scalable deployment through engineering, procurement, and construction (EPC) processes. This manuscript addresses this gap by proposing a phased development framework for AI data center expansion. The approach is designed to enable developers to meet aggressive time-to-market objectives while navigating multi-year constraints associated with interconnection approvals and lead times associated with the procurement of component equipment. A modular construction architecture is presented, along with a detailed analysis of integrated energy systems and the role of hybrid on-site generation in supporting incremental capacity growth. Electromagnetic transient simulations (EMT) are used to evaluate system performance, demonstrating that a combination of on-site natural gas generation and grid-forming energy storage can reliably support data center operations during early and intermediate deployment phases. The study further examines the transition to full grid interconnection, including the capability of the data center to operate in islanded mode during grid disturbances. Finally, the manuscript compares grid-forming control strategies for system reconnection and restoration under varying conditions.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers

Siqi Yan, Jiebao Zhang, Xi Yao, Juan Huang, Ye Shi

Published 2026-07-20 · arXiv · Credibility S

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing predictio…

Abstract, interpretation and reference

Abstract

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

The Cost and Network Limits of Space-Based AI Compute

Kees van Berkel

Published 2026-07-15 · arXiv · Credibility S

This paper evaluates whether large-scale AI data centers deployed in low-Earth orbit (LEO) could become a cost-effective alternative to terrestrial facilities. The analysis compares orbital and ground-based systems across launch cost, power generation, cooling, radiation exposure, and atmospheric reentry, as well as compute-network performance. A key distinction is the shift from terrestrial Clos networks to space-b…

Abstract, interpretation and reference

Abstract

This paper evaluates whether large-scale AI data centers deployed in low-Earth orbit (LEO) could become a cost-effective alternative to terrestrial facilities. The analysis compares orbital and ground-based systems across launch cost, power generation, cooling, radiation exposure, and atmospheric reentry, as well as compute-network performance. A key distinction is the shift from terrestrial Clos networks to space-based mesh networks using laser inter-satellite links. Using bisection bandwidth, bisection intensity, and roofline-style models, we show that while LEO-based inference may be feasible, training frontier-scale LLMs in orbit is unlikely to be competitive with terrestrial data centers.

Full text 中文海报
热管理与液冷 论文图示
Research Article芯片与算力

Hierarchical Multi-Agent Reinforcement Learning for Carbon-Aware AI Data Centers in Power Distribution Systems

Hyunsoo Lee, Panggah Prabawa, Dae-Hyun Choi, Joongheon Kim

Published 2026-07-03 · arXiv · Credibility S

Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy consumption-induced carbon emissions from AIDCs resulting from the rapid expansion of AI applications. This paper proposes a hierarchical carbon-aware multi-agent reinforcement learning (CA-MARL) framework for robust and efficient operations of AIDCs under uncertainties while ensur…

Abstract, interpretation and reference

Abstract

Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy consumption-induced carbon emissions from AIDCs resulting from the rapid expansion of AI applications. This paper proposes a hierarchical carbon-aware multi-agent reinforcement learning (CA-MARL) framework for robust and efficient operations of AIDCs under uncertainties while ensuring low-carbon operation of power distribution systems. The framework comprises a workload manager (WM) agent and multiple local AIDC agents trained using a multi-agent transformer method, corresponding to a global AIDC aggregator and a local AIDC operator, respectively. Leveraging AIDC operation data along with nodal carbon intensity (NCI) calculated from the carbon emission flow-integrated distribution system operator problem, the WM agent spatially allocates AI training and inference jobs among all AIDCs. Based on the jobs allocated from the WM agent and NCI information, each AIDC agent schedules economical and eco-friendly operations of the AIDC by performing the following tasks: i) temporal shifting of training jobs, ii) spatial allocation of training graphics processing unit (GPU) blocks and inference GPUs within the AIDC, and iii) control of the supply air temperature of the cooling system. The effectiveness of the proposed framework was assessed using an IEEE 33-node power distribution system.

Full text 中文海报
芯片与算力 论文图示
Research ArticleAI 运维优化

Storage as a Transmission Asset (SATA) for Large-Load Congestion Relief

Abanish Tiwari, Chandan Chaudhary, Yansong Pei, Mohammed Ben-Idris, Joydeep Mitra

Published 2026-07-05 · arXiv · Credibility S

Hyperscale data centers and other large concentrated loads can impose substantial new demand on existing transmission networks. If import corridors lack sufficient transfer capability, operators may need to curtail load, delay interconnection, or reinforce the network to maintain reliable service. An energy storage system (ESS) deployed as a storage-as-transmission asset (SATA) offers a non-wires alternative by prov…

Abstract, interpretation and reference

Abstract

Hyperscale data centers and other large concentrated loads can impose substantial new demand on existing transmission networks. If import corridors lack sufficient transfer capability, operators may need to curtail load, delay interconnection, or reinforce the network to maintain reliable service. An energy storage system (ESS) deployed as a storage-as-transmission asset (SATA) offers a non-wires alternative by providing operator-directed support to constrained import corridors. However, the operating-level reliability value of SATA dispatch remains insufficiently quantified. This paper evaluates operator-directed SATA using a day-ahead DC optimal power flow that co-optimizes generation, ESS dispatch, and load curtailment across Monte Carlo scenarios of demand and generator availability. Operating reliability is assessed using expected energy not served (EENS), loss-of-load hours (LOLH), and the conditional value at risk (CVaR) of daily unserved energy. Congestion-price and flow-sensitivity metrics are used to identify the limiting corridor and storage location. The interconnection is then screened to determine whether SATA is suitable, reinforcement is required, or storage would provide little transmission value. Results show that operator-directed SATA reduces average unserved energy, loss-of-load exposure, and tail risk compared with deploying the same ESS for pure arbitrage. These results demonstrate that the operating designation of storage is a primary driver of its transmission value.

Full text 中文海报
AI 运维优化 论文图示
Research Article算电协同

The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South

Muntaser Syed, Marius C. Silaghi, Sheikh Abujar, Sharun Akter Khushbu, Amal El Ahmad

Published 2026-07-15 · arXiv · Credibility S

Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four…

Abstract, interpretation and reference

Abstract

Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four cases: the United Arab Emirates, Bangladesh, India, and Africa (with a focus on Kenya). Using publicly available water stress data, grid carbon intensity factors, and GPU power specifications, we model the water consumption, energy demand, and carbon emissions of hypothetical sovereign AI deployments under multiple cooling technology scenarios. We find that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually in a country classified as ``extremely high'' water stress. In Bangladesh, sovereign AI policy documents call for centralized GPU procurement but do not address where to site data centers in a country where more than a fifth of the land floods in an average year and the power grid struggles to deliver reliable supply. We identify a sovereignty-sustainability trilemma in which no country can simultaneously maximize AI sovereignty, minimize environmental impact, and maintain affordable resource access for citizens. We propose design principles for environmentally responsible sovereign AI, including mandatory water usage effectiveness reporting, climate-vulnerability siting assessments, and a preference for frugal small language models over frontier pre-training in resource-constrained settings.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

A Hierarchical Semi-Markov Load Model for AI Data Centers Coupling Job Scheduling with Bulk-Synchronous-Parallel Power Dynamics

Chandan Chaudhary, Atri Bera, Cody Newlun, Mohammed Ben-Idris, Joydeep Mitra

Published 2026-07-13 · arXiv · Credibility S

AI data centers are emerging as a dominant new load class with their power dynamics fundamentally from conventional industrial loads. Inside a training job, the bulk-synchronous-parallel algorithm moves each node through compute, sync, and checkpoint steps, which swings power between full load and near idle within seconds. Across the whole facility, jobs arrive, take blocks of nodes for hours to days, then leave, so…

Abstract, interpretation and reference

Abstract

AI data centers are emerging as a dominant new load class with their power dynamics fundamentally from conventional industrial loads. Inside a training job, the bulk-synchronous-parallel algorithm moves each node through compute, sync, and checkpoint steps, which swings power between full load and near idle within seconds. Across the whole facility, jobs arrive, take blocks of nodes for hours to days, then leave, so the number of busy nodes changes daily, weekly, and yearly. This slower shift drives facility-wide swings and the peak demand that sets the size of the grid link. A model that looks only at within-job behavior, and treats the facility as a fixed set of busy nodes, smooths out these swings and misses the true peak-to-average ratio. This paper develops a hierarchical semi-Markov Data-Center (HSM-DC) load model that couples two layers across two timescales. A job-scheduling layer creates jobs through a non-homogeneous compound-Poisson process shaped by daily, weekly, and seasonal patterns, gives each job a heavy-tailed node count and length, and places jobs on a fixed pool of nodes on a first-come basis. A within-job layer moves each busy node through a five-state semi-Markov chain for the BSP steps, with state-based Ornstein-Uhlenbeck noise. Facility power comes from this changing node count and the per-node power, set to match measured node data and the facility's straight-line power-versus-load curve. Configured to the reference facility at the same scale, the model matches mean power, its spread, and the peak-to-average ratio across load levels, with fit scores of 0.9997, 0.92, and 0.82. It also matches the share of queued jobs to within one point at high load. Facility-wide swings and peak demand come from how jobs arrive and get scheduled, so grid planning must model that process, not just scale up a single node's power curve.

Full text 中文海报
算电协同 论文图示