智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 09-17

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article算电协同

From Grid to Chip: Power Architecture, Stability, and Flexibility of AI Data Centers

Yubo Song、Rui Kong、Takuro Umihara、Pooya Davari、Frede Blaabjerg、Subham Sahoo

Published 2026-09-10 · arXiv · Credibility S

The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a techno…

Abstract, interpretation and reference

Abstract

The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a technological perspective on AI data centers as grid-interactive computing systems. First, it reviews grid-integration bottlenecks, evolving connection policies, grid-code requirements, which has fostered new technological trends via spatio-temporal flexibility available through workload orchestration, cooling systems, on-site resources, and energy storage. Second, it maps the evolution of power-delivery architectures from medium-voltage grid interfaces to chip-level, discussing higher-voltage DC distribution, solid-state transformers, wide-bandgap devices, advanced chip-level power delivery, and liquid cooling. Third, it establishes a three-level stability framework spanning rack-level DC-bus dynamics, facility-level converter interactions, and system-level grid-coupled behavior. The framework connects dominant instability mechanisms, including constant power load effects, impedance interactions, forced oscillations, and operating-mode transitions, with suitable modeling, assessment, and mitigation approaches. Synthesizing these topics, this article highlights grid-to-chip co-design as a central requirement for scalable AI infrastructure, linking computing workloads, power-delivery systems, energy buffers, and grid operation.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用综述归纳和指标比较,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Yubo Song, Rui Kong, Takuro Umihara, 等. From Grid to Chip: Power Architecture, Stability, and Flexibility of AI Data Centers[J/OL]. (2026-09-10)[2026-09-17]. http://arxiv.org/abs/2609.11649v1.

Full text 中文海报
算电协同 论文图示
Research ArticleAI 运维优化

Quantifying AI data center nitrogen oxide (NO$_x$) emissions from space

Kevin D. Gauld、Daniel J. Varon、Nicholas Balasus、Daniel H. Cusworth

Published 2026-08-23 · arXiv · Credibility S

AI data center power demand is spurring rapid deployment of on- and near-site natural gas turbines. Nitrogen oxide (NO$_x$) pollution from this equipment is a growing concern but has not previously been quantified with atmospheric observations. Here we demonstrate space-based detection and quantification of NO$_x$ emissions from the SpaceXAI Colossus 2 power plant in Southaven, Mississippi. Using observations from t…

Abstract, interpretation and reference

Abstract

AI data center power demand is spurring rapid deployment of on- and near-site natural gas turbines. Nitrogen oxide (NO$_x$) pollution from this equipment is a growing concern but has not previously been quantified with atmospheric observations. Here we demonstrate space-based detection and quantification of NO$_x$ emissions from the SpaceXAI Colossus 2 power plant in Southaven, Mississippi. Using observations from the geostationary TEMPO satellite instrument, we detect a strong increase in local mean NO$_2$ column concentrations after the plant began operations in late 2025. We then use TEMPO to estimate two-week-average NO$_x$ source rates from August 2025 to mid-August 2026, calibrating against continuous emission monitoring system (CEMS) data from US power plants. TEMPO first detected NO$_x$ emissions in December 2025 at 460$\pm$180 kg h$^{-1}$. We find that emissions increased through August 2026, averaging 730$\pm$185 kg h$^{-1}$ after February 2026, roughly 16 times higher than expected from the facility's March 2026 permit for 41 turbines operating under best available control technology (BACT) requirements ($\sim$47 kg h$^{-1}$). Emissions at the expected level would be undetectable by our TEMPO analysis.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,AI 运维、负载预测和设施调优正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用仿真建模和情景分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断AI 工具是否能降低运维复杂度并提升可用性。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Kevin D. Gauld, Daniel J. Varon, Nicholas Balasus, 等. Quantifying AI data center nitrogen oxide (NO$_x$) emissions from space[J/OL]. (2026-08-23)[2026-09-17]. http://arxiv.org/abs/2608.22153v1.

Full text 中文海报
AI 运维优化 论文图示
Research Article算电协同

Grid-Mode-Aware Model Predictive Control of Hybrid Energy Storage Systems for AI Data Center Power Smoothing

Xin Chen

Published 2026-09-04 · arXiv · Credibility S

To facilitate the grid-friendly integration of highly variable AI data center loads, this paper proposes a grid-mode-aware model predictive control (G-MPC) framework for managing a hybrid energy storage system (HESS) to smooth grid-side power demand. The framework optimally coordinates a battery energy storage system (BESS) and a supercapacitor (SC) by solving a multi-step optimization problem in a receding-horizon …

Abstract, interpretation and reference

Abstract

To facilitate the grid-friendly integration of highly variable AI data center loads, this paper proposes a grid-mode-aware model predictive control (G-MPC) framework for managing a hybrid energy storage system (HESS) to smooth grid-side power demand. The framework optimally coordinates a battery energy storage system (BESS) and a supercapacitor (SC) by solving a multi-step optimization problem in a receding-horizon manner. In particular, band-pass filter dynamics are directly embedded in the G-MPC formulation to extract and suppress grid-side power components associated with vulnerable grid oscillatory modes, thus mitigating load-induced grid oscillations. The resulting G-MPC optimization jointly minimizes violations of grid-side power-envelope, ramp-rate, and modal-power requirements and the degradation and power-ramping costs of the BESS and SC, while satisfying power limits, state-of-charge limits, and other operational constraints. To enable real-time implementation, a fix-and-re-optimize algorithm is developed to solve each G-MPC problem efficiently while preventing simultaneous charging and discharging. Extensive simulations demonstrate the effectiveness, flexibility, and computational efficiency of the proposed framework. The results also highlight the importance of explicitly suppressing power components associated with vulnerable grid modes, rather than merely reducing overall load variations, to effectively mitigate grid oscillations.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Xin Chen. Grid-Mode-Aware Model Predictive Control of Hybrid Energy Storage Systems for AI Data Center Power Smoothing[J/OL]. (2026-09-04)[2026-09-17]. http://arxiv.org/abs/2609.04398v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Hosting Capacity Assessment of Data Centers with Voltage Ride-Through Capability in Power Systems

Pengyu Ren、Wei Sun、Fei Teng

Published 2026-09-03 · arXiv · Credibility S

Large data centers are emerging as concentrated, power-electronic grid loads whose abrupt disconnection or transfer to on-site backup supply during voltage disturbances can remove large demand from the power system, and may create a system-level stability problem. Their interconnection feasibility therefore depends not only on steady-state thermal and voltage limits, but also on whether internal power-conditioning s…

Abstract, interpretation and reference

Abstract

Large data centers are emerging as concentrated, power-electronic grid loads whose abrupt disconnection or transfer to on-site backup supply during voltage disturbances can remove large demand from the power system, and may create a system-level stability problem. Their interconnection feasibility therefore depends not only on steady-state thermal and voltage limits, but also on whether internal power-conditioning systems can maintain IT service while limiting customer-initiated load reduction. This paper presents a voltage ride-through (VRT)-aware data center and grid co-planning framework that couples transmission-level fault simulation with an internal data center ride-through model. Python-based dynamic simulations generate point-of-interconnection (POI) voltage trajectories under selected network faults, and the resulting waveforms drive an internal model incorporating IT and cooling-load dynamics, DC-link, Uninterruptible Power Supply (UPS) response, and converter apparent power limits. The IEEE 118-bus case study shows that internal VRT capability can become a binding interconnection constraint: steady-state planning alone can overestimate feasible data center capacity, whereas increased UPS converter headroom progressively restores hosting capacity. Under the reduced-order response models studied, the grid-forming mode provides greater ride-through margin than the current-limited grid-following mode under the same network fault conditions. The results further show that VRT constraints can materially change both the total hosting capacity of data centers and its spatial allocation across candidate interconnection buses.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Pengyu Ren, Wei Sun, Fei Teng. Hosting Capacity Assessment of Data Centers with Voltage Ride-Through Capability in Power Systems[J/OL]. (2026-09-03)[2026-09-17]. http://arxiv.org/abs/2609.03030v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Steady-State Equivalent Circuit Model for Data Center Loads

Muhammad Hamza Ali、Peng Sang、Hyeon Woo、Hyein Kang、Sungyun Choi、Amritanshu Pandey

Published 2026-08-18 · arXiv · Credibility S

Planners currently represent data centers as aggregate constant-PQ or ZIP loads in steady-state interconnection and contingency studies. These aggregate models are computationally convenient. However, they obscure the electrical relationship between computational workloads, server utilization, and grid-side demand. They ignore the internal power-electronic conversion stages of IT loads and assume homogeneous workloa…

Abstract, interpretation and reference

Abstract

Planners currently represent data centers as aggregate constant-PQ or ZIP loads in steady-state interconnection and contingency studies. These aggregate models are computationally convenient. However, they obscure the electrical relationship between computational workloads, server utilization, and grid-side demand. They ignore the internal power-electronic conversion stages of IT loads and assume homogeneous workload distributions across the compute clusters. This hides operating-point-dependent converter losses and efficiency variations. We propose a steady-state equivalent-circuit model (ECM) for data centers, which explicitly builds circuit models for IT loads, power supply units, cooling, and auxiliary systems. For power supply units, the equivalent circuit model explicitly represents internal power-electronic conversion stages. For IT loads, we develop a utilization-dependent server power model, and we combine it with loss-aware ECMs of power supply units. This approach captures the grid-side impact of heterogeneous workload distributions while preserving compatibility with conventional power-flow analysis. We evaluate this data center ECM in large-scale transmission power flows, using Monte Carlo simulations under heterogeneous and homogeneous cluster utilization. In comparison with the fixed-efficiency constant-PQ model, the ECM predicts that the most stressed line exceeds its thermal limit in about 30% of Monte Carlo samples. The results further show that homogeneous server utilization overstates line-loading variability by 17%-46% relative to heterogeneous server utilization, depending on the intra-cluster workload correlation.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Muhammad Hamza Ali, Peng Sang, Hyeon Woo, 等. Steady-State Equivalent Circuit Model for Data Center Loads[J/OL]. (2026-08-18)[2026-09-17]. http://arxiv.org/abs/2608.17925v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Flexible Training Workloads in Large-Scale AI Data Centers for Transient-Stability Support in Transmission-Constrained Power Systems

Jae-Kyeong Kim

Published 2026-08-31 · arXiv · Credibility S

The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support.…

Abstract, interpretation and reference

Abstract

The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support. To this end, this paper proposes training-induced load surge (TILS), a fast demand-side strategy that initiates or resumes flexible AI training workloads after fault clearing to increase active-power demand at electrically effective locations. The resulting load increase allows accelerating generators to supply additional electrical power, thereby reducing the accelerating-power imbalance and limiting the first-swing rotor-angle excursion. The underlying mechanism is first clarified in a single-machine infinite-bus (SMIB) system and then evaluated in the IEEE 39-bus system and a large-scale Korean power system. Results across all three systems demonstrate that TILS can increase the transient-stability-constrained generation limit. Larger responses, earlier activation, and siting at buses with a stronger electrical influence on the critical generators provide greater generation-limit increases. These results suggest that the upward load-response capability of AI data centers can provide complementary transient-stability support when sufficient electrical headroom, flexible workloads, and reliable grid-triggered activation are available.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用文献摘要中的模型、实验或案例分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Jae-Kyeong Kim. Flexible Training Workloads in Large-Scale AI Data Centers for Transient-Stability Support in Transmission-Constrained Power Systems[J/OL]. (2026-08-31)[2026-09-17]. http://arxiv.org/abs/2608.30901v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Exploiting the Benefits of V2B Application on Peak Shaving of Data Center Loads

Arya Joshi、Hamed Haggi、Chinmay Morankar

Published 2026-09-01 · arXiv · Credibility S

The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cool…

Abstract, interpretation and reference

Abstract

The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cooling loads. Initially, a literature review was conducted considering V2B constraints and optimization methods including SoC limitations, EV participation, tariffs, and building loads. This analysis was then used to develop a conceptual case study of a 10 MW data center in Loudoun County, VA by simulating a temperature-dependent load profile and adjusting the V2B participation of 40 commercial and passenger EVs. Simulation results indicate that, depending on seasonal variations in cooling load demands, strategic deployment of V2B assets between 12-5pm can offset gross cooling loads by 13-36%.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用综述归纳和指标比较,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Arya Joshi, Hamed Haggi, Chinmay Morankar. Exploiting the Benefits of V2B Application on Peak Shaving of Data Center Loads[J/OL]. (2026-09-01)[2026-09-17]. http://arxiv.org/abs/2609.00204v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Real-Time Edge-based Detection of Correlated AI Data-Center Load Episodes

Chandan Chaudhary、Abanish Tiwari、Yansong Pei、Mohammed Ben-Idris、Joydeep Mitra

Published 2026-08-24 · arXiv · Credibility S

Artificial-intelligence data centers running bulk-synchronous training can impose sub-second power swings. When several facilities synchronize their training cycles, these load variations become spatially correlated and amplify the aggregate disturbance on the grid. A grid operator without access to data-center telemetry must infer this correlation from electrical measurements alone. However, the required observatio…

Abstract, interpretation and reference

Abstract

Artificial-intelligence data centers running bulk-synchronous training can impose sub-second power swings. When several facilities synchronize their training cycles, these load variations become spatially correlated and amplify the aggregate disturbance on the grid. A grid operator without access to data-center telemetry must infer this correlation from electrical measurements alone. However, the required observation time and the feasibility of detection on substation-deployable hardware remain uncharacterized. This paper develops a correlation-based detection method to classify the multi-facility operating regime from cross-facility power measurements. Analytical derivations and experimental validation show that the resulting detection confidence increases with the observation-window length at a rate governed by the load correlation time. The method is demonstrated in a real-time hardware-in-the-loop testbed, where load setpoints generated from a validated semi-Markov data-center load model are applied to an electromagnetic-transient grid simulation on a Real-Time Digital Simulator. A compact classifier built on pairwise power correlations runs on an edge device in this loop and determines whether the data-center load variations are independent or spatially correlated. The cross-facility correlation separates the independent and correlated cases across independent realizations. The held-out detection accuracy improves with the observation window, consistent with the predicted relation. A raw-waveform network fails to generalize, supporting pairwise correlation as the discriminative signal. The detector executes in real time on commodity edge hardware. A closed-loop demonstration against the running simulator tracks a regime change within one observation window.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用实验验证、原型测试或测量对比,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Chandan Chaudhary, Abanish Tiwari, Yansong Pei, 等. Real-Time Edge-based Detection of Correlated AI Data-Center Load Episodes[J/OL]. (2026-08-24)[2026-09-17]. http://arxiv.org/abs/2608.22719v1.

Full text 中文海报
算电协同 论文图示