智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 07-30

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article算电协同

Emission-Forecasting-Based Spatial-Temporal Carbon Response: A Multi-Agent Attention-Enhanced Deep Learning Framework

Feiyu Cai, Jing Qiu, Yi Yang, Chenxi Zhang, Xinlei Wang, Baichuan Liu, Junhua Zhao

Published 2026-07-29 · arXiv · Credibility S

As a major contributor to carbon emissions, the decarbonization of power systems has garnered significant societal attention. Nodal carbon intensity (NCI), a critical factor in carbon-oriented demand response, has traditionally been determined through ex-post calculations. However, this ex-post approach introduces latency in low-carbon dispatch. To address this, this paper presents a proactive ex-ante spatial-tempor…

Abstract, interpretation and reference

Abstract

As a major contributor to carbon emissions, the decarbonization of power systems has garnered significant societal attention. Nodal carbon intensity (NCI), a critical factor in carbon-oriented demand response, has traditionally been determined through ex-post calculations. However, this ex-post approach introduces latency in low-carbon dispatch. To address this, this paper presents a proactive ex-ante spatial-temporal carbon response framework. At its core, we develop a novel deep learning-based hierarchical design, enhanced by a dual-stage attention mechanism and a large language model (LLM)-based multi-agent cooperation system, to accurately forecast day-ahead NCI. This design effectively mitigates the impact of renewable energy uncertainty and enhances predictive resilience. On the demand side, the framework proposes a spatial-temporal carbon scheduling model that integrates geographically dispatchable loads (GDLs), including mobile energy storage systems (MESSs) and distributed data centers (DDCs). Leveraging high-accuracy day-ahead NCI predictions, the framework can effectively reduce system emissions by quickly responding to carbon intensity fluctuations. The proposed framework is tested on the modified IEEE 33-bus system. According to the simulation results, the impacts of proposed framework on dispatching latency and emission outcomes are analyzed. The results demonstrate that under a one-hour reduction in carbon scheduling latency, the proposed model and methodology can achieve over 30% emission reduction. This research breaks through the limitations of passive carbon accounting, advancing toward proactive carbon management. It offers an intelligent solution that accelerates the transition to cleaner power systems while directly supporting sustainable production goals.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Planning Waste-to-Energy-Coupled AI Data Centers Through Grade-Matched Cooling and Corridor Screening

Qi He, Chunyu Qu, Wenjie Zuo

Published 2026-07-28 · arXiv · Credibility S

AI data-center growth is increasingly constrained by limited deliverable electricity, interconnection capacity, and cooling demand. This study develops a boundary-consistent screening framework for waste-to-energy (WtE)-coupled AI data-center cooling. It treats cooling as an energy service that can be supplied through grade matching rather than only through electricity-driven mechanical chilling. The framework trans…

Abstract, interpretation and reference

Abstract

AI data-center growth is increasingly constrained by limited deliverable electricity, interconnection capacity, and cooling demand. This study develops a boundary-consistent screening framework for waste-to-energy (WtE)-coupled AI data-center cooling. It treats cooling as an energy service that can be supplied through grade matching rather than only through electricity-driven mechanical chilling. The framework translates plant-side exportable heat into corridor-level planning metrics by accounting for thermal attenuation, absorption conversion, and parasitic electricity for delivery and auxiliaries. In a reference case, a regulated WtE plant processing 1500 t/day of municipal solid waste at 10 MJ/kg provides about 78.1 MWth of exportable heat. At a 20 km corridor, this yields about 53.0 MW of delivered cooling and 8.0 MWe of net avoided cooling electricity after parasitic loads. The coupled system is governed by operating regimes rather than a single efficiency score. Under baseline assumptions, full thermal coverage extends to about 20.9 km, the quality-adjusted criterion remains positive to about 22.9 km, and net electricity relief remains positive to about 44.7 km. For a 1 GW IT campus at 70 percent utilization and a 5 km corridor, net grid relief ranges from about 116.9 to 264.4 MW across scenarios. The required WtE footprint ranges from about 3 to 148 representative plants, or 0.6 to 40 full-load-equivalent plants at a 25 percent displacement target. The framework identifies when WtE-coupled cooling is corridor-feasible, when hybrid operation is required, and when infrastructure scale becomes the binding constraint. It is intended for screening and comparison, not project-specific hydraulic or plant-cycle design.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

From Individual to Shared Ownership: A Coalitional Game Approach to Sustainable Co-investment

暂无可靠最新数据

Published 2026-07-29 · arXiv · Credibility S

This paper proposes a cooperative game-theoretic framework for sustainable co-investment in shared infrastructure under regulatory incentives. Multiple heterogeneous operators co-invest in a common infrastructure whose production capability evolves over time and is subject to operational variability. A regulator supports the deployment through incentive mechanisms designed to align individual economic investment obj…

Abstract, interpretation and reference

Abstract

This paper proposes a cooperative game-theoretic framework for sustainable co-investment in shared infrastructure under regulatory incentives. Multiple heterogeneous operators co-invest in a common infrastructure whose production capability evolves over time and is subject to operational variability. A regulator supports the deployment through incentive mechanisms designed to align individual economic investment objectives with the coalitional one. We formulate the co-investment problem as a transferable-utility (TU) coalitional game in which the value generated by cooperation depends on heterogeneous operational profiles, dynamic resource availability, investment costs, and regulatory incentive level. We show that the proposed coalitional game can be reformulated as a linear production game (LPG), whose dual prices yield a constructive and stable allocation of the cooperative surplus. Finally, we illustrate the proposed framework through a case study on co-investment among data center operators in shared renewable energy infrastructure, supported by government subsidies promoting renewable energy consumption.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers

Siqi Yan, Jiebao Zhang, Xi Yao, Juan Huang, Ye Shi

Published 2026-07-20 · arXiv · Credibility S

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing predictio…

Abstract, interpretation and reference

Abstract

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures

Soham Ghosh, Nabil Mohammed, Mohammad Ashraf Hossain Sadi

Published 2026-07-19 · arXiv · Credibility S

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operati…

Abstract, interpretation and reference

Abstract

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operational impacts of AI data centers, as well as their potential role as grid-interactive assets, limited attention has been given to the challenges associated with their scalable deployment through engineering, procurement, and construction (EPC) processes. This manuscript addresses this gap by proposing a phased development framework for AI data center expansion. The approach is designed to enable developers to meet aggressive time-to-market objectives while navigating multi-year constraints associated with interconnection approvals and lead times associated with the procurement of component equipment. A modular construction architecture is presented, along with a detailed analysis of integrated energy systems and the role of hybrid on-site generation in supporting incremental capacity growth. Electromagnetic transient simulations (EMT) are used to evaluate system performance, demonstrating that a combination of on-site natural gas generation and grid-forming energy storage can reliably support data center operations during early and intermediate deployment phases. The study further examines the transition to full grid interconnection, including the capability of the data center to operate in islanded mode during grid disturbances. Finally, the manuscript compares grid-forming control strategies for system reconnection and restoration under varying conditions.

Full text 中文海报
算电协同 论文图示
Research Article芯片与算力

Hierarchical Multi-Agent Reinforcement Learning for Carbon-Aware AI Data Centers in Power Distribution Systems

Hyunsoo Lee, Panggah Prabawa, Dae-Hyun Choi, Joongheon Kim

Published 2026-07-03 · arXiv · Credibility S

Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy consumption-induced carbon emissions from AIDCs resulting from the rapid expansion of AI applications. This paper proposes a hierarchical carbon-aware multi-agent reinforcement learning (CA-MARL) framework for robust and efficient operations of AIDCs under uncertainties while ensur…

Abstract, interpretation and reference

Abstract

Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy consumption-induced carbon emissions from AIDCs resulting from the rapid expansion of AI applications. This paper proposes a hierarchical carbon-aware multi-agent reinforcement learning (CA-MARL) framework for robust and efficient operations of AIDCs under uncertainties while ensuring low-carbon operation of power distribution systems. The framework comprises a workload manager (WM) agent and multiple local AIDC agents trained using a multi-agent transformer method, corresponding to a global AIDC aggregator and a local AIDC operator, respectively. Leveraging AIDC operation data along with nodal carbon intensity (NCI) calculated from the carbon emission flow-integrated distribution system operator problem, the WM agent spatially allocates AI training and inference jobs among all AIDCs. Based on the jobs allocated from the WM agent and NCI information, each AIDC agent schedules economical and eco-friendly operations of the AIDC by performing the following tasks: i) temporal shifting of training jobs, ii) spatial allocation of training graphics processing unit (GPU) blocks and inference GPUs within the AIDC, and iii) control of the supply air temperature of the cooling system. The effectiveness of the proposed framework was assessed using an IEEE 33-node power distribution system.

Full text 中文海报
芯片与算力 论文图示
Research Article算电协同

How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization

Xinyi Wu, Siyuan Liu, Ali Jadbabaie

Published 2026-07-08 · arXiv · Credibility S

Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly. We study what determines this frequency usage and propose a data-centered explanation: RoPE frequencies are selected to match the relative-distance structure of the training data. Viewing each frequency as a positional lens, we formalize a field-resolution…

Abstract, interpretation and reference

Abstract

Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly. We study what determines this frequency usage and propose a data-centered explanation: RoPE frequencies are selected to match the relative-distance structure of the training data. Viewing each frequency as a positional lens, we formalize a field-resolution tradeoff and show that, for a data-induced dependency profile of width $W$, the optimal frequency scales as $1/W$. This frequency-matching principle explains controlled observations on synthetic and text-based data, and suggests that the mid-low frequency bands observed in language models arise from the multi-scale dependency structure of natural language. We further connect frequency selection to position-interpolation-based length generalization: scaling frequencies down expands the effective field while reducing resolution. This helps when longer-context dependencies are approximate dilations of those seen during training, but can fail when relevant dependencies do not scale with context length. Empirically, we show that natural language exhibits approximate self-similarity across positional scales, explaining why test-time frequency scaling can support long-context generalization. Overall, our results identify a data-driven mechanism behind emergent RoPE frequency usage and show that long-context generalization depends on two forms of scale matching: between learned frequencies and training-time dependencies, and between frequency scaling and how those dependencies extend to longer contexts.

Full text 中文海报
算电协同 论文图示
Research ArticleAI 运维优化

Storage as a Transmission Asset (SATA) for Large-Load Congestion Relief

Abanish Tiwari, Chandan Chaudhary, Yansong Pei, Mohammed Ben-Idris, Joydeep Mitra

Published 2026-07-05 · arXiv · Credibility S

Hyperscale data centers and other large concentrated loads can impose substantial new demand on existing transmission networks. If import corridors lack sufficient transfer capability, operators may need to curtail load, delay interconnection, or reinforce the network to maintain reliable service. An energy storage system (ESS) deployed as a storage-as-transmission asset (SATA) offers a non-wires alternative by prov…

Abstract, interpretation and reference

Abstract

Hyperscale data centers and other large concentrated loads can impose substantial new demand on existing transmission networks. If import corridors lack sufficient transfer capability, operators may need to curtail load, delay interconnection, or reinforce the network to maintain reliable service. An energy storage system (ESS) deployed as a storage-as-transmission asset (SATA) offers a non-wires alternative by providing operator-directed support to constrained import corridors. However, the operating-level reliability value of SATA dispatch remains insufficiently quantified. This paper evaluates operator-directed SATA using a day-ahead DC optimal power flow that co-optimizes generation, ESS dispatch, and load curtailment across Monte Carlo scenarios of demand and generator availability. Operating reliability is assessed using expected energy not served (EENS), loss-of-load hours (LOLH), and the conditional value at risk (CVaR) of daily unserved energy. Congestion-price and flow-sensitivity metrics are used to identify the limiting corridor and storage location. The interconnection is then screened to determine whether SATA is suitable, reinforcement is required, or storage would provide little transmission value. Results show that operator-directed SATA reduces average unserved energy, loss-of-load exposure, and tail risk compared with deploying the same ESS for pure arbitrage. These results demonstrate that the operating designation of storage is a primary driver of its transmission value.

Full text 中文海报
AI 运维优化 论文图示