智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 09-24

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article算电协同

Privacy-Preserving Coordinated Operation of Power Grids and AI Data Centers: A Checkpoint-Aware Three-Phase Scheme

Ziang Liu、Ruizhang Yang、Xin Cui、Francis Yunhe Hou

Published 2026-09-22 · arXiv · Credibility S

The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase …

Abstract, interpretation and reference

Abstract

The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase transmission congestion risks. Coordinating AIDC operation with grid scheduling under these unique operational characteristics is challenging because grid and AIDC operators are generally unwilling to share proprietary data and decision-making authority. This paper proposes a hierarchical privacy-preserving coordinated operation scheme between the power grid and AIDCs to address this gap. The proposed scheme contains three phases. In Phase I, the grid operator computes a certified inner approximation of the AIDCs security region for subsequent coordination. In Phase II, the AIDC operator coordinates training and inference AIDCs to optimize workload allocation within the certified security region and generate power schedules and checkpoint alerts. In Phase III, the grid operator solves a checkpoint-aware two-stage robust optimal power flow (OPF) considering renewable generation and checkpoint uncertainties. By exchanging only compact interface information, the framework preserves the privacy of both grid and AIDCs, avoids frequent iterative communication, and enables secure coordination with guaranteed feasibility. Numerical studies on a modified IEEE 14-bus system and a modified NYISO system demonstrate the effectiveness, robustness, and security of the proposed framework.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Ziang Liu, Ruizhang Yang, Xin Cui, 等. Privacy-Preserving Coordinated Operation of Power Grids and AI Data Centers: A Checkpoint-Aware Three-Phase Scheme[J/OL]. (2026-09-22)[2026-09-24]. http://arxiv.org/abs/2609.26365v1.

Full text 中文海报
算电协同 论文图示
Research Article能效优化

Beyond PUE: A Local Impact Audit Framework for Data Center Environmental Accountability

Sharifa Sultana、Syed Ishtiaque Ahmed

Published 2026-09-20 · arXiv · Credibility S

Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of …

Abstract, interpretation and reference

Abstract

Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of 2026 alone, local opposition delayed or canceled roughly $130 billion in projects across the United States, driven overwhelmingly by recurring concerns over water use, power demand, infrastructure capacity, and transparency rather than internal efficiency, matching the total for all of 2025 [11]. We propose a five-category local impact audit framework covering efficiency, water stewardship, carbon and renewables, regulatory compliance, and local disclosure. The framework is designed for recurring quarterly assessment and independent verification against public records. We illustrate its application using publicly available data from three Illinois facilities that are currently at the center of local policy disputes, and we examine the data-access barriers that constrain independent verification. We position this framework as both a research contribution and a practical instrument for county-level policymakers evaluating data center permitting and moratorium decisions.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,PUE/WUE、能效指标和运营成本控制正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断不同能效指标是否真实反映节能和成本收益。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Sharifa Sultana, Syed Ishtiaque Ahmed. Beyond PUE: A Local Impact Audit Framework for Data Center Environmental Accountability[J/OL]. (2026-09-20)[2026-09-24]. http://arxiv.org/abs/2609.23421v1.

Full text 中文海报
能效优化 论文图示
Research Article芯片与算力

ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters

Rui Lu、Rui Ge、Huanghuang Liang、Xiaobo Zhou、Dan Wang

Published 2026-09-14 · arXiv · Credibility S

Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU…

Abstract, interpretation and reference

Abstract

Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU-plus-cooling energy while satisfying thermal safety and latency SLO constraints. We present ETCInfer, an energy-efficient, thermal-aware scheduler that selects a pre-job Computer Room Air Conditioner (CRAC) setpoint and adapts per-GPU frequency and micro-batch size during execution. ETCInfer builds compact physics-informed control models by calibrating GPU heat generation, chassis heat dissipation, CRAC power, and prefill/decode latency relations from telemetry. These models estimate hidden thermal states and time-to-throttle, enabling the scheduler to evaluate energy, temperature, and latency before applying an action. We formulate this joint setpoint--frequency--micro-batch control problem as a partially observable Markov decision process and design ETCAdapter, a learning-based controller that minimizes per-job energy under thermal safety and SLO constraints. We implement ETCInfer as a coordination layer over typical inference and cluster management stacks. Evaluation across real-trace simulation and validation experiments shows that ETCInfer reduces total job energy by up to 33.1%, thermal throttle exposure by up to 92.9%, and keeps SLO violation rates below 0.7% even at ambient temperatures up to $48^{\circ}\mathrm{C}$.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,芯片、服务器和高密度算力部署正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用实验验证、原型测试或测量对比,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向算力硬件、边缘计算或模型部署对基础设施的牵引。意义:对日报读者而言,它可用于判断芯片路线和服务器密度变化如何传导到机房设计。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Rui Lu, Rui Ge, Huanghuang Liang, 等. ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters[J/OL]. (2026-09-14)[2026-09-24]. http://arxiv.org/abs/2609.15230v1.

Full text 中文海报
芯片与算力 论文图示
Research Article热管理与液冷

Could Underwater Data Centers Pose a Risk to AI Treaty Verification?

James Teague、Ashmita Rajmohan、Yannick Muehlhaeuser

Published 2026-09-16 · arXiv · Credibility S

Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate cons…

Abstract, interpretation and reference

Abstract

Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate construction and maintenance complexity relative to land-based facilities, and analyse the feasibility of a 100,000 H100-equivalent training run underwater. We find that power delivery and cooling are tractable, but interconnect and the hands-on maintenance that large training runs require are severe obstacles - surmountable only by a well-resourced state actor accepting large cost and schedule penalties, and only where concealment, rather than efficiency, is the objective. We then assess detectability through thermal, acoustic, optical and synthetic-aperture-radar (SAR) surveillance. Thermal detection of an operational pod is unlikely outside shallow, calm water; acoustic detection is marginally more effective, but faces limitations in attribution; and optical/SAR monitoring is most powerful during construction and maintenance, when the pressure-vessel fabrication base and the cable-laying fleet create distinctive signatures for AIS-tracking. We conclude that UDCs are a comparatively unlikely evasion route relative to underground or industrially disguised land-based facilities, but the residual risk is non-zero and warrants operationalising the detection modalities discussed.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,液冷、热管理和数据中心能效正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用文献摘要中的模型、实验或案例分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断液冷方案、热管理路线和高密度部署节奏。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

James Teague, Ashmita Rajmohan, Yannick Muehlhaeuser. Could Underwater Data Centers Pose a Risk to AI Treaty Verification?[J/OL]. (2026-09-16)[2026-09-24]. http://arxiv.org/abs/2609.18824v1.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

CATS: A Carbon-Aware Task Simulator for Reducing AI Data Center Emissions

Dayuan Chen、Ziliang Zong

Published 2026-09-14 · arXiv · Credibility S

The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper. First, we compile a global alignment dataset unifying 140 operationa…

Abstract, interpretation and reference

Abstract

The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper. First, we compile a global alignment dataset unifying 140 operational and planned cloud regions across 8 major providers with five-minute carbon-intensity traces for 145 grid regions from 2022 to 2024, revealing that 50% of current sites lie in medium-to-high carbon-intensity grids, indicating a siting-carbon mismatch and unrealized carbon reduction potential. Second, we develop CATS (Carbon-Aware Task Simulator), a flexible trace-driven framework that profiles six AI inference tasks across multiple GPU types, synthesizes realistic diurnal curve, geographical and task mixes, and SLA constraints, and evaluates spatial and temporal schedulers against two baselines while reporting comprehensive metrics including carbon emissions, energy consumption, runtime, queue delay, and hardware utilization. Third, we quantify achievable CO2 savings under realistic constraints: in a 24-hour trace with 600,000 tasks at fleet utilization of 0.37, spatial shifting reduces CO2 by 38.4% versus speed-first baseline, while temporal shifting yields 16% savings with bounded SLA violations at 3.27%. These results advocate locating future data centers in low carbon-intensity grids and demonstrate that carbon-aware scheduling on today's fleets can achieve substantial operational emissions reduction.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Dayuan Chen, Ziliang Zong. CATS: A Carbon-Aware Task Simulator for Reducing AI Data Center Emissions[J/OL]. (2026-09-14)[2026-09-24]. http://arxiv.org/abs/2609.14775v1.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

Convective Heat Transfer Optimization for Liquid Cooling Plates Driven by Field Synergy and Fractal Geometry

Zixu Han、Peng Zhang

Published 2026-09-11 · arXiv · Credibility S

The rapid development of liquid-cooled data centers has imposed imperative demands on the performance of liquid cooling plate. The density-based topology optimization (TO) is an effective approach to resolving the growing thermal-hydraulic performance requirements of liquid cooling plate. However, existing TO methods can hardly optimize convective heat transfer directly which is the intrinsic heat transfer mechanism…

Abstract, interpretation and reference

Abstract

The rapid development of liquid-cooled data centers has imposed imperative demands on the performance of liquid cooling plate. The density-based topology optimization (TO) is an effective approach to resolving the growing thermal-hydraulic performance requirements of liquid cooling plate. However, existing TO methods can hardly optimize convective heat transfer directly which is the intrinsic heat transfer mechanism, due to the highly complex and evolving structural topologies, varying flow and temperature fields, making it extremely challenging to explicitly describe the heat transfer coefficient and heat transfer area during TO process. A convective heat transfer topology optimization (CTO) method is proposed in this study, where the iteratively evolving heat transfer coefficient is explicitly depicted by the field synergy theory in the thermal objective, and directly described by the velocity and temperature fields without relying on specific geometry. Combined with the explicit depiction of heat transfer area by the fractal geometry theory, a CTO framework is built for a direct optimization of convective heat transfer under both the laminar and turbulent flow conditions. The CTO tends to generate more hierarchical and directional structural topologies in optimization results, which is conducive to reducing low-velocity stagnation zones and improving flow direction in branched channels, achieving enhanced synergy and thermal-hydraulic performance in the optimized liquid cooling plates. Compared with the TO results without incorporation of field synergy theory, the CTO can reduce average temperature rise by 20% while improving the Nusselt number by 15% under laminar flow conditions, and reduce maximum temperature rise by 10.2% and pressure drop by 25% under turbulent flow conditions.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,液冷、热管理和数据中心能效正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断液冷方案、热管理路线和高密度部署节奏。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Zixu Han, Peng Zhang. Convective Heat Transfer Optimization for Liquid Cooling Plates Driven by Field Synergy and Fractal Geometry[J/OL]. (2026-09-11)[2026-09-24]. http://arxiv.org/abs/2609.12344v1.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

Spatial LLM Workload Shifting Needs Foresight: Model Commitment for AI Data Center Operation under Power Grid Constraints

Bojun Du、Hongyang Jia、Tonghui Li、Qingchun Hou、Ze Wang、Ershun Du、Ning Zhang

Published 2026-09-09 · arXiv · Credibility S

AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter prop…

Abstract, interpretation and reference

Abstract

AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter proposes model commitment (MC), a mixed-integer linear programming framework that jointly schedules model deployment and cross-site request routing under power constraints and electricity-price signals. First, MC formulates the intertemporal coupling introduced by model replica loading. Second, it translates prefill and decode latency requirements into the amount of demand that each replica can serve. Case studies based on real-world data show that MC enables AI data center operators to achieve a 100% service rate under time-varying grid conditions and reduce total operating cost by 29.0%.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Bojun Du, Hongyang Jia, Tonghui Li, 等. Spatial LLM Workload Shifting Needs Foresight: Model Commitment for AI Data Center Operation under Power Grid Constraints[J/OL]. (2026-09-09)[2026-09-24]. http://arxiv.org/abs/2609.09787v1.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

From Grid to Chip: Power Architecture, Stability, and Flexibility of AI Data Centers

Yubo Song、Rui Kong、Takuro Umihara、Pooya Davari、Frede Blaabjerg、Subham Sahoo

Published 2026-09-10 · arXiv · Credibility S

The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a techno…

Abstract, interpretation and reference

Abstract

The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a technological perspective on AI data centers as grid-interactive computing systems. First, it reviews grid-integration bottlenecks, evolving connection policies, grid-code requirements, which has fostered new technological trends via spatio-temporal flexibility available through workload orchestration, cooling systems, on-site resources, and energy storage. Second, it maps the evolution of power-delivery architectures from medium-voltage grid interfaces to chip-level, discussing higher-voltage DC distribution, solid-state transformers, wide-bandgap devices, advanced chip-level power delivery, and liquid cooling. Third, it establishes a three-level stability framework spanning rack-level DC-bus dynamics, facility-level converter interactions, and system-level grid-coupled behavior. The framework connects dominant instability mechanisms, including constant power load effects, impedance interactions, forced oscillations, and operating-mode transitions, with suitable modeling, assessment, and mitigation approaches. Synthesizing these topics, this article highlights grid-to-chip co-design as a central requirement for scalable AI infrastructure, linking computing workloads, power-delivery systems, energy buffers, and grid operation.

中文解读

背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用综述归纳和指标比较,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Yubo Song, Rui Kong, Takuro Umihara, 等. From Grid to Chip: Power Architecture, Stability, and Flexibility of AI Data Centers[J/OL]. (2026-09-10)[2026-09-24]. http://arxiv.org/abs/2609.11649v1.

Full text 中文海报
算电协同 论文图示