智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 10-01

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article芯片与算力

Integrated Thermal and Power Management for Wave-Powered Subsea Data Centers via Nonlinear Model Predictive Control

Wanqun Yang、Jun Chen

Published 2026-09-28 · arXiv · Credibility S

This paper develops an integrated modeling and nonlinear model predictive control (NMPC) framework for coordinating thermal management, flexible workload scheduling, wave-power utilization, and battery operation in a wave-powered subsea data center. Realistic data center workloads are constructed from job-level CPU, memory, and GPU measurements from the MIT Supercloud dataset and divided into interactive and delay-t…

Abstract, interpretation and reference

Abstract

This paper develops an integrated modeling and nonlinear model predictive control (NMPC) framework for coordinating thermal management, flexible workload scheduling, wave-power utilization, and battery operation in a wave-powered subsea data center. Realistic data center workloads are constructed from job-level CPU, memory, and GPU measurements from the MIT Supercloud dataset and divided into interactive and delay-tolerant flexible jobs. Thermal behavior is represented by a three-node lumped model of the IT equipment, recirculating nitrogen, and pressure hull with surrounding seawater as the thermal boundary. The NMPC jointly optimizes the flexible workload power budget and cooling command subject to thermal, battery, and workload constraints. Closed-loop simulations under different workload, thermal, battery, and renewable-generation conditions demonstrate that the proposed framework maintains thermal safety while adapting cooling operation and flexible workload execution to wave-power availability and battery state-of-charge. The parametric studies show that battery capacity and wave-generation capacity strongly affect battery availability and flexible-workload queue accumulation, while excessive renewable generation capacity may lead to increased energy curtailment. Monte Carlo and distance-correlation analyses further show that flexible-job delay is relatively insensitive to the investigated system parameters, whereas terminal battery state-of-charge is primarily influenced by battery energy capacity and wave generation capacity.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,芯片、服务器和高密度算力部署正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断芯片路线和服务器密度变化如何传导到机房设计。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Wanqun Yang, Jun Chen. Integrated Thermal and Power Management for Wave-Powered Subsea Data Centers via Nonlinear Model Predictive Control[J/OL]. (2026-09-28)[2026-10-01]. http://arxiv.org/abs/2609.34109v1.

Full text 中文海报
芯片与算力 论文图示
Research Article算电协同

Grid-Forming E-STATCOMs for Stable Integration of Large-Scale Data Centers: Modeling and Control

Prabhat Ranjan Bana、Novan Zakkia、Jean-Philippe Hasler、Christer Danielsson

Published 2026-09-27 · arXiv · Credibility S

The rapid expansion of large-scale AI data centers (AIDC) is introducing new stability challenges, particularly in weak or low-inertia networks characterized by fast, step-like demand variations and strict requirements on voltage and dynamic performance. This paper investigates the use of grid-forming (GFM) Enhanced STATCOMs (E-STATCOMs) to support reliable integration of such facilities. A power-admittance-based li…

Abstract, interpretation and reference

Abstract

The rapid expansion of large-scale AI data centers (AIDC) is introducing new stability challenges, particularly in weak or low-inertia networks characterized by fast, step-like demand variations and strict requirements on voltage and dynamic performance. This paper investigates the use of grid-forming (GFM) Enhanced STATCOMs (E-STATCOMs) to support reliable integration of such facilities. A power-admittance-based linear modelling framework is developed to capture system interactions and is validated through detailed EMT simulations. The results demonstrate that E-STATCOMs provide fast, well-damped responses to abrupt load changes while effectively mitigating low-frequency oscillations and interactions with network resonances. By enabling tunable dynamic behavior via a load balancer, virtual impedance, and coordinated active-reactive power support, the proposed approach allows precise shaping of system response and improved regulation at the point of connection. These features make E-STATCOMs a flexible and scalable solution for integrating large data centers into weak grids and long transmission systems, supported by a design-oriented framework that facilitates parameter selection and performance assessment without extensive reliance on EMT studies to meet grid codes and AIDC interconnection requirements.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Prabhat Ranjan Bana, Novan Zakkia, Jean-Philippe Hasler, 等. Grid-Forming E-STATCOMs for Stable Integration of Large-Scale Data Centers: Modeling and Control[J/OL]. (2026-09-27)[2026-10-01]. http://arxiv.org/abs/2609.33553v1.

Full text 中文海报
算电协同 论文图示
Research Article能效优化

Beyond PUE: A Local Impact Audit Framework for Data Center Environmental Accountability

Sharifa Sultana、Syed Ishtiaque Ahmed

Published 2026-09-20 · arXiv · Credibility S

Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of …

Abstract, interpretation and reference

Abstract

Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of 2026 alone, local opposition delayed or canceled roughly $130 billion in projects across the United States, driven overwhelmingly by recurring concerns over water use, power demand, infrastructure capacity, and transparency rather than internal efficiency, matching the total for all of 2025 [11]. We propose a five-category local impact audit framework covering efficiency, water stewardship, carbon and renewables, regulatory compliance, and local disclosure. The framework is designed for recurring quarterly assessment and independent verification against public records. We illustrate its application using publicly available data from three Illinois facilities that are currently at the center of local policy disputes, and we examine the data-access barriers that constrain independent verification. We position this framework as both a research contribution and a practical instrument for county-level policymakers evaluating data center permitting and moratorium decisions.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,PUE/WUE、能效指标和运营成本控制正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断不同能效指标是否真实反映节能和成本收益。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Sharifa Sultana, Syed Ishtiaque Ahmed. Beyond PUE: A Local Impact Audit Framework for Data Center Environmental Accountability[J/OL]. (2026-09-20)[2026-10-01]. http://arxiv.org/abs/2609.23421v1.

Full text 中文海报
能效优化 论文图示
Research Article热管理与液冷

Could Underwater Data Centers Pose a Risk to AI Treaty Verification?

James Teague、Ashmita Rajmohan、Yannick Muehlhaeuser

Published 2026-09-16 · arXiv · Credibility S

Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate cons…

Abstract, interpretation and reference

Abstract

Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate construction and maintenance complexity relative to land-based facilities, and analyse the feasibility of a 100,000 H100-equivalent training run underwater. We find that power delivery and cooling are tractable, but interconnect and the hands-on maintenance that large training runs require are severe obstacles - surmountable only by a well-resourced state actor accepting large cost and schedule penalties, and only where concealment, rather than efficiency, is the objective. We then assess detectability through thermal, acoustic, optical and synthetic-aperture-radar (SAR) surveillance. Thermal detection of an operational pod is unlikely outside shallow, calm water; acoustic detection is marginally more effective, but faces limitations in attribution; and optical/SAR monitoring is most powerful during construction and maintenance, when the pressure-vessel fabrication base and the cable-laying fleet create distinctive signatures for AIS-tracking. We conclude that UDCs are a comparatively unlikely evasion route relative to underground or industrially disguised land-based facilities, but the residual risk is non-zero and warrants operationalising the detection modalities discussed.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,液冷、热管理和数据中心能效正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用文献摘要中的模型、实验或案例分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断液冷方案、热管理路线和高密度部署节奏。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

James Teague, Ashmita Rajmohan, Yannick Muehlhaeuser. Could Underwater Data Centers Pose a Risk to AI Treaty Verification?[J/OL]. (2026-09-16)[2026-10-01]. http://arxiv.org/abs/2609.18824v1.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

Job Class Thermal Intent Aware Liquid Cooling Allocation for AI Data Centers

Krishna Chaitanya Sunkara

Published 2026-09-25 · arXiv · Credibility S

GPU-dense AI data centers need to run on liquid cooling as air simply cannot shed the heat at these power densities. Yet the cooling loops themselves are blind to what workloads are about to run; they crank up flow only after a sensor catches a temperature climb, which can take 30 to 50 seconds. We built Job-Class Thermal Intent (JCTI) to close that window. The scheduler already knows a job is coming and what class …

Abstract, interpretation and reference

Abstract

GPU-dense AI data centers need to run on liquid cooling as air simply cannot shed the heat at these power densities. Yet the cooling loops themselves are blind to what workloads are about to run; they crank up flow only after a sensor catches a temperature climb, which can take 30 to 50 seconds. We built Job-Class Thermal Intent (JCTI) to close that window. The scheduler already knows a job is coming and what class it belongs to; JCTI feeds that information straight to the cooling controller so it can stage coolant before the heat shows up. We pulled the thermal signatures for each job class out of MLPerf GPU power traces and tuned arrival patterns against Alibaba cluster data. Over 120 paired Monte Carlo trials the numbers come out to 56.4% fewer thermal violations and 60.2% less cumulative overshoot than a straight PI loop. As AI data centers evolving towards gigawatt grid loads with highly fluctuating power swings, thermally-aware scheduling reduces sudden demand and improves load prediction in grid side. Cooling and scheduling have been running as two separate systems for years despite each one knowing something the other needs, JCTI wires them together.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Krishna Chaitanya Sunkara. Job Class Thermal Intent Aware Liquid Cooling Allocation for AI Data Centers[J/OL]. (2026-09-25)[2026-10-01]. http://arxiv.org/abs/2609.30785v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI

Farnaz Farid、Tashfia Towkee、Sania Nasreen、Sami bin Azad

Published 2026-09-25 · arXiv · Credibility S

As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center cooling and electricity generation receive little attention. To address this gap, we introduce SustainAI, a water-aware, closed-loop framework incorporating env…

Abstract, interpretation and reference

Abstract

As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center cooling and electricity generation receive little attention. To address this gap, we introduce SustainAI, a water-aware, closed-loop framework incorporating environmental accountability into AI deployment. SustainAI integrates real-time water metering, a hallucination-aware penalty model, and a water-aware routing algorithm that accounts for regional water stress. Evaluated via Small Language Models (SLMs) extracting health misinformation, results reveal an 11-fold variation in water footprint across geographically distributed data centers (0.0477 mL to 0.5360 mL per inference). Across 1,335 inference runs, the system consumed approximately 399 mL of water but produced only 240 correct outputs, demonstrating that substantial resources are spent on inaccurate responses. Crucially, SustainAI extends beyond technical optimization through a Care by Design lens, framing AI sustainability around relational ethics, regional equity, and ecological stewardship. By combining water monitoring, adaptive accountability, and Care by Design principles, SustainAI provides a practical foundation for integrating ethical care and environmental responsibility into AI infrastructure design and lifecycle management.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Farnaz Farid, Tashfia Towkee, Sania Nasreen, 等. Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI[J/OL]. (2026-09-25)[2026-10-01]. http://arxiv.org/abs/2609.30747v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Data center cooling choices shift water impacts across the grid: An integrated water-energy model for sustainable data center development

Garrett Alston、Nancy Love、Rabab Haider

Published 2026-09-22 · arXiv · Credibility S

Data centers are being developed at an unprecedented pace, yet their energy and water impacts, and the spatial and temporal distribution of these impacts, remain poorly characterized. Data centers consume water for cooling (direct) and through electricity generation (indirect). Decisions on siting and cooling technology result in water-energy trade-offs that extend impacts beyond the facility's location. Existing as…

Abstract, interpretation and reference

Abstract

Data centers are being developed at an unprecedented pace, yet their energy and water impacts, and the spatial and temporal distribution of these impacts, remain poorly characterized. Data centers consume water for cooling (direct) and through electricity generation (indirect). Decisions on siting and cooling technology result in water-energy trade-offs that extend impacts beyond the facility's location. Existing assessment frameworks rely on facility efficiency metrics and average grid water intensity factors, suppressing the temporal impacts of data center load and generation availability. They also attribute indirect consumption to the facility's location rather than to the generators (and corresponding hydrologic regions) that respond to the added load, misattributing spatial impacts. To close this gap, we develop a computational model of the data center-energy-water nexus that links facility cooling and electricity demand with hourly economic dispatch, generator-level water consumption, and monthly subbasin depletion. Built on open-source data, the model resolves where and when water is consumed, and where this consumption compounds existing water risk or creates new risk. Using the model, we study different cooling configurations and proposed developments in the state of Michigan. Air-cooled data centers halve total water consumption relative to evaporative cooling, but increase electricity demand and raise indirect water consumption by one-third, shifting the water footprint from the facility to generators. Mapping these changes to subbasins reveals depletion increases beyond the data center sites, in regions that facility-level reporting may overlook. These results show that data center water and energy impacts cannot be assessed in isolation, motivating the need for integrated modeling to inform siting, design, and reporting practices.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Garrett Alston, Nancy Love, Rabab Haider. Data center cooling choices shift water impacts across the grid: An integrated water-energy model for sustainable data center development[J/OL]. (2026-09-22)[2026-10-01]. http://arxiv.org/abs/2609.25437v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Privacy-Preserving Coordinated Operation of Power Grids and AI Data Centers: A Checkpoint-Aware Three-Phase Scheme

Ziang Liu、Ruizhang Yang、Xin Cui、Francis Yunhe Hou

Published 2026-09-22 · arXiv · Credibility S

The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase …

Abstract, interpretation and reference

Abstract

The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase transmission congestion risks. Coordinating AIDC operation with grid scheduling under these unique operational characteristics is challenging because grid and AIDC operators are generally unwilling to share proprietary data and decision-making authority. This paper proposes a hierarchical privacy-preserving coordinated operation scheme between the power grid and AIDCs to address this gap. The proposed scheme contains three phases. In Phase I, the grid operator computes a certified inner approximation of the AIDCs security region for subsequent coordination. In Phase II, the AIDC operator coordinates training and inference AIDCs to optimize workload allocation within the certified security region and generate power schedules and checkpoint alerts. In Phase III, the grid operator solves a checkpoint-aware two-stage robust optimal power flow (OPF) considering renewable generation and checkpoint uncertainties. By exchanging only compact interface information, the framework preserves the privacy of both grid and AIDCs, avoids frequent iterative communication, and enables secure coordination with guaranteed feasibility. Numerical studies on a modified IEEE 14-bus system and a modified NYISO system demonstrate the effectiveness, robustness, and security of the proposed framework.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Ziang Liu, Ruizhang Yang, Xin Cui, 等. Privacy-Preserving Coordinated Operation of Power Grids and AI Data Centers: A Checkpoint-Aware Three-Phase Scheme[J/OL]. (2026-09-22)[2026-10-01]. http://arxiv.org/abs/2609.26365v1.

Full text 中文海报
算电协同 论文图示