SimBIG: Field-level Simulation-Based Inference of Galaxy Clustering

📄 arXiv: 2310.15256v1 📥 PDF

作者: Pablo Lemos, Liam Parker, ChangHoon Hahn, Shirley Ho, Michael Eickenberg, Jiamin Hou, Elena Massara, Chirag Modi, Azadeh Moradinezhad Dizgah, Bruno Regaldo-Saint Blancard, David Spergel

分类: astro-ph.CO, cs.LG

发布日期: 2023-10-23

备注: 14 pages, 4 figures. A previous version of the paper was published in the ICML 2023 Workshop on Machine Learning for Astrophysics


💡 一句话要点

提出SimBIG框架以解决宇宙学参数推断问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control)

关键词: 宇宙学 星系聚类 模拟推断 正则化流 卷积神经网络 数据压缩 非高斯特征

📋 核心要点

  1. 现有的星系聚类分析方法主要依赖于总结统计量,未能充分利用星系分布的非线性和非高斯特征。
  2. 论文提出SimBIG框架,通过正则化流进行模拟推断,利用卷积神经网络对星系场进行数据压缩,提升参数推断精度。
  3. 实验结果显示,Ω_m的约束与传统方法一致,而σ_8的约束精度提高了2.65倍,同时还推断出哈勃常数H_0。

📝 摘要(中文)

本文首次基于场级分析的模拟推断(SBI)宇宙学参数,针对标准的星系聚类分析方法的不足,提出了SimBIG前向建模框架,利用正则化流进行推断。通过对BOSS CMASS星系样本的应用,使用卷积神经网络进行数据压缩,推断出Ω_m和σ_8的约束,后者的约束精度比传统方法提高了2.65倍。此外,本文还从星系聚类中推断出哈勃常数H_0。该研究展示了在不同前向模型下的鲁棒性,为未来的星系调查提供了新的方法论。

🔬 方法详解

问题定义:本文旨在解决传统星系聚类分析方法未能充分利用星系分布的非线性和非高斯特征的问题。现有方法主要依赖于总结统计量,如功率谱,导致信息损失。

核心思路:论文提出的SimBIG框架通过正则化流进行模拟推断,能够更全面地捕捉星系分布的复杂特征,从而提高宇宙学参数的推断精度。

技术框架:整体架构包括数据预处理、卷积神经网络模型构建、正则化流推断和参数约束四个主要模块。首先对星系数据进行压缩,然后通过正则化流进行参数推断。

关键创新:最重要的技术创新在于引入了正则化流作为推断工具,使得能够利用更多的非高斯信息,从而显著提高了参数约束的精度。

关键设计:在网络结构上,采用了卷积神经网络,并结合随机权重平均技术进行数据压缩,损失函数设计上则注重于优化推断精度和鲁棒性。具体参数设置和网络结构细节在实验部分进行了详细描述。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,本文对Ω_m的约束为0.267^{+0.033}{-0.029},与传统方法一致,而σ_8的约束为0.762^{+0.036}{-0.035},精度提高了2.65倍。此外,哈勃常数H_0的推断结果为64.5 ± 3.8 km/s/Mpc,显示出更强的约束能力。

🎯 应用场景

该研究的潜在应用领域包括未来的大规模星系调查,如DESI、PFS和Euclid等。通过引入SimBIG框架,研究人员可以更有效地从星系聚类数据中提取宇宙学信息,推动宇宙学研究的进展。

📄 摘要(原文)

We present the first simulation-based inference (SBI) of cosmological parameters from field-level analysis of galaxy clustering. Standard galaxy clustering analyses rely on analyzing summary statistics, such as the power spectrum, $P_\ell$, with analytic models based on perturbation theory. Consequently, they do not fully exploit the non-linear and non-Gaussian features of the galaxy distribution. To address these limitations, we use the {\sc SimBIG} forward modelling framework to perform SBI using normalizing flows. We apply SimBIG to a subset of the BOSS CMASS galaxy sample using a convolutional neural network with stochastic weight averaging to perform massive data compression of the galaxy field. We infer constraints on $Ω_m = 0.267^{+0.033}{-0.029}$ and $σ_8=0.762^{+0.036}{-0.035}$. While our constraints on $Ω_m$ are in-line with standard $P_\ell$ analyses, those on $σ_8$ are $2.65\times$ tighter. Our analysis also provides constraints on the Hubble constant $H_0=64.5 \pm 3.8 \ {\rm km / s / Mpc}$ from galaxy clustering alone. This higher constraining power comes from additional non-Gaussian cosmological information, inaccessible with $P_\ell$. We demonstrate the robustness of our analysis by showcasing our ability to infer unbiased cosmological constraints from a series of test simulations that are constructed using different forward models than the one used in our training dataset. This work not only presents competitive cosmological constraints but also introduces novel methods for leveraging additional cosmological information in upcoming galaxy surveys like DESI, PFS, and Euclid.