Research Article热管理与液冷
Seung Hyeon Ham、Yu Min Kim、In Hyeok Choi、Jeong Woo Han
Published 2026-08-26 · arXiv · Credibility S
The rapid growth of generative AI has intensified the need for efficient heat dissipation in large-scale data centers. To control heat flow, thermal metamaterials with layered structures have been widely used, which impart the anisotropic properties of thermal conductivities. However, the conventional effective medium approximation (EMA) often fails to provide accurate predictions in systems with a high thermal cond…
Abstract, interpretation and reference
Abstract
The rapid growth of generative AI has intensified the need for efficient heat dissipation in large-scale data centers. To control heat flow, thermal metamaterials with layered structures have been widely used, which impart the anisotropic properties of thermal conductivities. However, the conventional effective medium approximation (EMA) often fails to provide accurate predictions in systems with a high thermal conductivity contrast between adjacent layers embedded in a background medium. Here, we generalize the EMA by introducing two corrective coefficients that extend its validity to regimes where the conventional EMA was previously inapplicable, i.e., high-contrast thermal metamaterials with the background medium. Notably, one of these coefficients that we proposed has the same mathematical form as the Fresnel reflection coefficient in optics. This allows us to interpret the "reflection-like" behavior of heat flow as it penetrates adjacent layers with high thermal contrast. Our findings suggest that heat diffusion, traditionally viewed as a purely dissipative process, can be understood intuitively through the framework of ray optics.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,液冷、热管理和数据中心能效正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断液冷方案、热管理路线和高密度部署节奏。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Seung Hyeon Ham, Yu Min Kim, In Hyeok Choi, 等. Generalizing Thermal Transport in High-Contrast Metamaterials through Interfacial Fresnel Reflection[J/OL]. (2026-08-26)[2026-09-22]. http://arxiv.org/abs/2608.25499v1.
Research Article算电协同
Chandan Chaudhary、Abanish Tiwari、Yansong Pei、Mohammed Ben-Idris、Joydeep Mitra
Published 2026-08-24 · arXiv · Credibility S
Artificial-intelligence data centers running bulk-synchronous training can impose sub-second power swings. When several facilities synchronize their training cycles, these load variations become spatially correlated and amplify the aggregate disturbance on the grid. A grid operator without access to data-center telemetry must infer this correlation from electrical measurements alone. However, the required observatio…
Abstract, interpretation and reference
Abstract
Artificial-intelligence data centers running bulk-synchronous training can impose sub-second power swings. When several facilities synchronize their training cycles, these load variations become spatially correlated and amplify the aggregate disturbance on the grid. A grid operator without access to data-center telemetry must infer this correlation from electrical measurements alone. However, the required observation time and the feasibility of detection on substation-deployable hardware remain uncharacterized. This paper develops a correlation-based detection method to classify the multi-facility operating regime from cross-facility power measurements. Analytical derivations and experimental validation show that the resulting detection confidence increases with the observation-window length at a rate governed by the load correlation time. The method is demonstrated in a real-time hardware-in-the-loop testbed, where load setpoints generated from a validated semi-Markov data-center load model are applied to an electromagnetic-transient grid simulation on a Real-Time Digital Simulator. A compact classifier built on pairwise power correlations runs on an edge device in this loop and determines whether the data-center load variations are independent or spatially correlated. The cross-facility correlation separates the independent and correlated cases across independent realizations. The held-out detection accuracy improves with the observation window, consistent with the predicted relation. A raw-waveform network fails to generalize, supporting pairwise correlation as the discriminative signal. The detector executes in real time on commodity edge hardware. A closed-loop demonstration against the running simulator tracks a regime change within one observation window.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用实验验证、原型测试或测量对比,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Chandan Chaudhary, Abanish Tiwari, Yansong Pei, 等. Real-Time Edge-based Detection of Correlated AI Data-Center Load Episodes[J/OL]. (2026-08-24)[2026-09-22]. http://arxiv.org/abs/2608.22719v1.
Research ArticleAI 运维优化
Kevin D. Gauld、Daniel J. Varon、Nicholas Balasus、Daniel H. Cusworth
Published 2026-08-23 · arXiv · Credibility S
AI data center power demand is spurring rapid deployment of on- and near-site natural gas turbines. Nitrogen oxide (NO$_x$) pollution from this equipment is a growing concern but has not previously been quantified with atmospheric observations. Here we demonstrate space-based detection and quantification of NO$_x$ emissions from the SpaceXAI Colossus 2 power plant in Southaven, Mississippi. Using observations from t…
Abstract, interpretation and reference
Abstract
AI data center power demand is spurring rapid deployment of on- and near-site natural gas turbines. Nitrogen oxide (NO$_x$) pollution from this equipment is a growing concern but has not previously been quantified with atmospheric observations. Here we demonstrate space-based detection and quantification of NO$_x$ emissions from the SpaceXAI Colossus 2 power plant in Southaven, Mississippi. Using observations from the geostationary TEMPO satellite instrument, we detect a strong increase in local mean NO$_2$ column concentrations after the plant began operations in late 2025. We then use TEMPO to estimate two-week-average NO$_x$ source rates from August 2025 to mid-August 2026, calibrating against continuous emission monitoring system (CEMS) data from US power plants. TEMPO first detected NO$_x$ emissions in December 2025 at 460$\pm$180 kg h$^{-1}$. We find that emissions increased through August 2026, averaging 730$\pm$185 kg h$^{-1}$ after February 2026, roughly 16 times higher than expected from the facility's March 2026 permit for 41 turbines operating under best available control technology (BACT) requirements ($\sim$47 kg h$^{-1}$). Emissions at the expected level would be undetectable by our TEMPO analysis.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,AI 运维、负载预测和设施调优正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用仿真建模和情景分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断AI 工具是否能降低运维复杂度并提升可用性。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Kevin D. Gauld, Daniel J. Varon, Nicholas Balasus, 等. Quantifying AI data center nitrogen oxide (NO$_x$) emissions from space[J/OL]. (2026-08-23)[2026-09-22]. http://arxiv.org/abs/2608.22153v2.
Research Article能效优化
Sharifa Sultana、Syed Ishtiaque Ahmed
Published 2026-09-20 · arXiv · Credibility S
Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of …
Abstract, interpretation and reference
Abstract
Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of 2026 alone, local opposition delayed or canceled roughly $130 billion in projects across the United States, driven overwhelmingly by recurring concerns over water use, power demand, infrastructure capacity, and transparency rather than internal efficiency, matching the total for all of 2025 [11]. We propose a five-category local impact audit framework covering efficiency, water stewardship, carbon and renewables, regulatory compliance, and local disclosure. The framework is designed for recurring quarterly assessment and independent verification against public records. We illustrate its application using publicly available data from three Illinois facilities that are currently at the center of local policy disputes, and we examine the data-access barriers that constrain independent verification. We position this framework as both a research contribution and a practical instrument for county-level policymakers evaluating data center permitting and moratorium decisions.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,PUE/WUE、能效指标和运营成本控制正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断不同能效指标是否真实反映节能和成本收益。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Sharifa Sultana, Syed Ishtiaque Ahmed. Beyond PUE: A Local Impact Audit Framework for Data Center Environmental Accountability[J/OL]. (2026-09-20)[2026-09-22]. http://arxiv.org/abs/2609.23421v1.
Research Article算电协同
Arya Joshi、Hamed Haggi、Chinmay Morankar
Published 2026-09-01 · arXiv · Credibility S
The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cool…
Abstract, interpretation and reference
Abstract
The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cooling loads. Initially, a literature review was conducted considering V2B constraints and optimization methods including SoC limitations, EV participation, tariffs, and building loads. This analysis was then used to develop a conceptual case study of a 10 MW data center in Loudoun County, VA by simulating a temperature-dependent load profile and adjusting the V2B participation of 40 commercial and passenger EVs. Simulation results indicate that, depending on seasonal variations in cooling load demands, strategic deployment of V2B assets between 12-5pm can offset gross cooling loads by 13-36%.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用综述归纳和指标比较,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Arya Joshi, Hamed Haggi, Chinmay Morankar. Exploiting the Benefits of V2B Application on Peak Shaving of Data Center Loads[J/OL]. (2026-09-01)[2026-09-22]. http://arxiv.org/abs/2609.00204v1.
Research ArticleAI 运维优化
Slava G. Turyshev
Published 2026-08-27 · arXiv · Credibility S
Megawatt-class orbital data centers require continuous maintenance, replacement, inventory, and service capacity in addition to spacecraft power/thermal systems. We formulate an analytical lifecycle framework for permanent/transient failures, modular orbital replacement units, robotic servicing, spare inventory, scheduled technology refresh, correlated faults, cybersecurity, optional human support. The model combine…
Abstract, interpretation and reference
Abstract
Megawatt-class orbital data centers require continuous maintenance, replacement, inventory, and service capacity in addition to spacecraft power/thermal systems. We formulate an analytical lifecycle framework for permanent/transient failures, modular orbital replacement units, robotic servicing, spare inventory, scheduled technology refresh, correlated faults, cybersecurity, optional human support. The model combines nonhomogeneous component hazards, capacity-weighted availability, multiclass robotic-service capacity, Poisson base-stock inventory, replacement-flow accounting, human-support break-even relations. For a 1 MW cluster with 10 active 100 kW nodes, 1 reserve node, ~200 5 kW compute cartridges, low, nominal, high deployed-mass allocations span ~50-75 kg/kW. Assumptions yield 70.2 random or life-limited interventions and 323-349 planned refresh operations/(MW-year), for a total of 393-419 standardized operations/(MW-year). Analysis gives a first-generation logistics of 5.3-9.0 t/(MW-year), with a nominal case of ~ 6.6 t/(MW year), 560-700 productive robot-hors/(MW-year). Planned refresh exceeds random replacement under the stated component populations, hazards, 3-15-year intervals. At ~400 standardized operations/(MW-year), the post-internal-recovery exception probability is <$10^{-3}$, with an objective near $10^{-4}$ at large scale; terminal non-recovery $p_U$ requires a smaller mission-level allocation. The target catastrophic-loss hazard for a 100 kW node is 0.01-0.03 1/yr. Parametric workload and cost cases place contingency visits at 10s of MWs, periodic campaigns at 10-100s of MWs, dedicated personnel at several 100 MWs to GWs. The reference first deployment is uncrewed, autonomously fault-managed, robotically maintainable, supported by specific inventory based on a 6-month replenishment horizon, compatible with later human access without permanent habitation.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,AI 运维、负载预测和设施调优正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断AI 工具是否能降低运维复杂度并提升可用性。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Slava G. Turyshev. Operations, Maintenance, and Industrial Scaling of MW-Class Orbital Data Centers[J/OL]. (2026-08-27)[2026-09-22]. http://arxiv.org/abs/2608.27499v1.
Research Article算电协同
Jae-Kyeong Kim
Published 2026-08-31 · arXiv · Credibility S
The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support.…
Abstract, interpretation and reference
Abstract
The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support. To this end, this paper proposes training-induced load surge (TILS), a fast demand-side strategy that initiates or resumes flexible AI training workloads after fault clearing to increase active-power demand at electrically effective locations. The resulting load increase allows accelerating generators to supply additional electrical power, thereby reducing the accelerating-power imbalance and limiting the first-swing rotor-angle excursion. The underlying mechanism is first clarified in a single-machine infinite-bus (SMIB) system and then evaluated in the IEEE 39-bus system and a large-scale Korean power system. Results across all three systems demonstrate that TILS can increase the transient-stability-constrained generation limit. Larger responses, earlier activation, and siting at buses with a stronger electrical influence on the critical generators provide greater generation-limit increases. These results suggest that the upward load-response capability of AI data centers can provide complementary transient-stability support when sufficient electrical headroom, flexible workloads, and reliable grid-triggered activation are available.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用文献摘要中的模型、实验或案例分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Jae-Kyeong Kim. Flexible Training Workloads in Large-Scale AI Data Centers for Transient-Stability Support in Transmission-Constrained Power Systems[J/OL]. (2026-08-31)[2026-09-22]. http://arxiv.org/abs/2608.30901v1.
Research Article算电协同
Hassan Zahid Butt、Rida Fatima、Xingpeng Li
Published 2026-08-30 · arXiv · Credibility S
Securing grid interconnection capacity has become a bottleneck for AI data center projects and can take longer than constructing the facilities themselves. This mismatch can delay deployment for years, making early interconnection planning essential. This paper develops ICP-AI, an interconnection capacity planning framework from a data center developer's perspective. The framework minimizes grid import capacity unde…
Abstract, interpretation and reference
Abstract
Securing grid interconnection capacity has become a bottleneck for AI data center projects and can take longer than constructing the facilities themselves. This mismatch can delay deployment for years, making early interconnection planning essential. This paper develops ICP-AI, an interconnection capacity planning framework from a data center developer's perspective. The framework minimizes grid import capacity under a prescribed onsite investment budget while jointly sizing photovoltaic (PV) and battery energy storage system (BESS) resources and scheduling deadline constrained workload flexibility. A secondary refinement fixes the minimum grid capacity and selects the minimum-investment PV-BESS portfolio among solutions that achieve that capacity. The framework is evaluated using monthly composite stress profiles across varying temporal assumptions, load shapes, flexible load fractions, and deferral windows. Results show that interconnection capacity reduction depends strongly on the planning environment: at a $100M budget, it is about 6% for the high load factor baseline, exceeds 10% under monthly average solar availability, and reaches 13.3% for a more diurnal load. At a $10M budget, 5% flexible load with a 1 h workload deferral window reduces BESS capacity from 15.30 to 4.87 MWh while increasing capacity reduction from 4.43% to 4.84%. To test sensitivity to temporal compression, the model is also solved over the full 8,760 h chronology, which preserves the main capacity and flexibility trends. Overall, ICP-AI quantifies the interconnection capacity and infrastructure substitution value of workload flexibility, providing an investment-interconnection frontier to support capital allocation and early project planning in constrained grid environments.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Hassan Zahid Butt, Rida Fatima, Xingpeng Li. Minimizing Grid Interconnection Capacity Requirements for AI Data Centers: A Developer-Side Planning Framework with Onsite Resources and Workload Flexibility[J/OL]. (2026-08-30)[2026-09-22]. http://arxiv.org/abs/2608.29359v1.