AeroMELD: A Linear Embedding of Aerosol Populations for Diagnostics and Latent Dynamics
作者: Ehsan Saleh, Saba Ghaffari, Wenhan Tang, Jeffrey H. Curtis, Lekha Patel, Peter A. Bosler, Nicole Riemer, Matthew West
分类: cs.LG, physics.ao-ph
发布日期: 2026-07-13
备注: 34 pages, 12 figures
💡 一句话要点
提出AeroMELD框架以改进气溶胶群体的表征与动态学习
🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)
关键词: 气溶胶表征 机器学习 潜在动态 深度学习 气候模型 环境监测
📋 核心要点
- 现有的气溶胶表征方法由于结构假设的限制,无法有效捕捉气溶胶的成分多样性和混合状态。
- AeroMELD框架通过构建低维潜在变量,保留气溶胶群体的数学结构,解决了传统自编码器的不足。
- 实验结果显示,AeroMELD能够准确重构气溶胶的质量和数量分布,提升了对气溶胶过程的理解和模拟能力。
📝 摘要(中文)
准确表征大气气溶胶群体对于模拟气溶胶-云相互作用、辐射强迫和冰核化至关重要。然而,现有的简化方案由于结构假设的限制,无法充分捕捉成分多样性和混合状态。机器学习方法提供了更灵活的表征,但标准自编码器无法保持气溶胶群体的数学结构。本文提出了AeroMELD(气溶胶测量嵌入用于潜在动态),一个数学基础框架,用于构建保留该结构的低维潜在变量。实验表明,该框架能够准确重构气溶胶分布、光学系数和冰冻行为,同时保留所需的线性群体结构,为混合机器学习-物理模型奠定了基础。
🔬 方法详解
问题定义:本文旨在解决现有气溶胶表征方法无法有效捕捉气溶胶成分多样性和混合状态的问题。现有的简化方案由于结构假设的限制,导致表征能力不足。
核心思路:AeroMELD框架的核心思想是构建一个数学基础的低维潜在变量,能够保留气溶胶群体的数学结构,从而支持物理上有意义的过程操作。
技术框架:该框架包括一个线性编码器,采用尺度-形状分解,明确表示总数浓度,并通过粒子嵌入的重心组合给出潜在形状。聚合的潜在状态保留了深度集模型的诊断表达能力,同时将非线性后聚合阶段移动到学习的诊断映射中。
关键创新:AeroMELD的主要创新在于其能够在潜在空间中直接学习气溶胶过程的演变,保留了线性群体结构,区别于传统自编码器的非线性特性。
关键设计:该框架通过直接编码加权粒子群体,而不是分箱气溶胶状态,使用粒子解析数据作为真实值,确保了潜在空间的准确重构。
🖼️ 关键图片
📊 实验亮点
实验结果表明,AeroMELD能够准确重构气溶胶的质量和数量分布、CCN光谱及光学系数,且在保持线性群体结构的同时,显著提升了对气溶胶过程的模拟能力,展示了其在混合机器学习-物理模型中的应用潜力。
🎯 应用场景
AeroMELD框架在气候模型、环境监测和大气科学等领域具有广泛的应用潜力。通过准确表征气溶胶群体,该研究能够帮助科学家更好地理解气溶胶对气候变化的影响,并为相关政策制定提供科学依据。
📄 摘要(原文)
Accurately representing atmospheric aerosol populations is essential for simulating aerosol-cloud interactions, radiative forcing, and ice nucleation, yet existing reduced schemes impose structural assumptions that limit their ability to capture composition diversity and mixing state. Machine-learning approaches offer more flexible representations, but standard autoencoders do not preserve the mathematical structure of aerosol populations and therefore cannot support physically meaningful process operators. We introduce AeroMELD (Aerosol Measure Embedding for Latent Dynamics), a mathematically grounded framework for constructing low-dimensional latent variables that retain this structure. We show that any permutation-invariant linear encoder must take a scale-shape decomposition, with total number concentration represented explicitly and latent shape given by a barycentric combination of per-particle embeddings. This aggregated latent state retains the diagnostic expressiveness of a Deep Sets model by moving the nonlinear post-aggregation stage into the learned diagnostic map while preserving latent linearity. Using particle-resolved data as ground truth, we encode weighted particle populations directly rather than binned aerosol states; size-resolved mass and number distributions serve only as diagnostic targets and visual summaries. The latent space accurately reconstructs these distributions, CCN spectra, optical coefficients, and immersion-freezing behavior while preserving the linear population structure needed for hybrid ML-physics models. Although the experiments focus on diagnostic reconstruction, the embedding is designed so that emissions and mixing can be represented exactly and nonlinear microphysical processes learned in a controlled latent space. This work establishes a foundation for learning aerosol-process evolution directly in latent space.