Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
作者: Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li, Boxuan Zhao, Yonghong Tian, Bartosz A. Grzybowski, Fanyang Mo
分类: physics.chem-ph, cs.LG, physics.comp-ph, physics.data-an
发布日期: 2026-07-22
备注: 6 figures
💡 一句话要点
提出假设-精炼学习以解决有机结构解析问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 有机结构解析 假设-精炼学习 多模态光谱 分子生成 数据驱动方法
📋 核心要点
- 核心问题:从光谱数据中解析分子结构面临逆问题的欠定性,现有方法难以充分利用光谱信息。
- 方法要点:提出假设-精炼学习框架,结合光谱证据与大规模分子先验,构建QM9SPIN数据集和SpectroMol模型。
- 实验或效果:系统在模拟基准上达到93.8%准确率,并能有效适应实验光谱,展示出良好的数据驱动结构解析能力。
📝 摘要(中文)
从光谱数据中确定分子结构仍然是一个基本挑战,因为逆问题本质上是欠定的:单个光谱稀疏、低维,并且仅编码相对于可能分子空间的部分结构证据。我们通过将自动结构阐明形式化为可扩展的假设-精炼范式来解决这一挑战,该范式紧密结合光谱证据与大规模分子先验。为提供结构解析的NMR信号,我们构建了QM9SPIN数据集,包含多样的1D和2D光谱。基于此,我们引入了SpectroMol模型,该模型根据多模态光谱输入提出化学有效的分子假设。此外,我们开发了MS-Mol2Mol,一个高分辨率的质量约束分子生成器,确保全球成分一致性和化学上合理的精炼。该系统在模拟基准上实现了93.8%的顶级准确率,并有效适应从模拟到实验光谱的转变。
🔬 方法详解
问题定义:本论文旨在解决从光谱数据中解析有机分子结构的逆问题。现有方法由于光谱数据的稀疏性和低维性,往往无法充分利用光谱信息,导致解析结果不准确。
核心思路:论文提出了一种假设-精炼学习的框架,通过将光谱证据与大规模分子先验结合,形成一个可扩展的自动结构阐明方法。这种设计能够更好地利用现有的光谱数据,提高结构解析的准确性。
技术框架:整体架构包括两个主要模块:SpectroMol模型用于根据多模态光谱输入生成化学有效的分子假设,MS-Mol2Mol则是一个高分辨率的质量约束分子生成器,确保生成分子的化学合理性和成分一致性。
关键创新:最重要的技术创新在于将假设-精炼学习与大规模分子先验结合,形成了一种新的数据驱动的结构解析方法。这与传统方法的单一光谱分析形成了本质区别。
关键设计:在模型设计中,使用了多模态输入,结合了分子式、精确质量和不饱和度等信息,确保生成的分子在化学上合理。同时,训练过程中采用了特定的损失函数,以优化生成分子的质量和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,集成系统在模拟基准上达到了93.8%的顶级准确率,且在从模拟光谱到实验光谱的转变中表现出良好的适应性。通过质量引导的精炼,进一步提升了实验预测的准确性,展示了该方法的有效性。
🎯 应用场景
该研究的潜在应用领域包括药物发现、材料科学和化学合成等。通过自动化的结构解析方法,可以大幅提高分子设计和筛选的效率,推动新材料和新药物的开发,具有重要的实际价值和未来影响。
📄 摘要(原文)
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.