Entanglement as a Structural Complexity Axis: A PAC-Bayesian View of Generalization in Quantum Policies and Value Functions
作者: Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
分类: quant-ph, cs.LG
发布日期: 2026-07-07
💡 一句话要点
提出PAC-Bayesian框架以揭示量子策略的泛化特性
🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)
关键词: 量子强化学习 PAC-Bayesian 参数化量子电路 泛化能力 Fisher几何 纠缠连接性 量子策略设计
📋 核心要点
- 现有量子策略的泛化能力机制不明确,传统方法主要关注参数数量,忽视了有效维度的影响。
- 提出PAC-Bayesian框架,强调Fisher几何有效维度与纠缠的关系,作为量子策略设计的新视角。
- 实验结果显示,纠缠电路的泛化性能普遍低于非纠缠电路,且在不同任务中均表现出一致性。
📝 摘要(中文)
参数化量子电路(PQCs)在量子强化学习中被广泛应用于策略和价值函数,但其泛化能力的机制尚不明确。本文提出了一种PAC-Bayesian视角,认为泛化能力受电路引发的Fisher几何有效维度的影响,而非仅仅是电路参数的数量。研究发现,纠缠连接性作为复杂性的独立轴,影响了电路的有效维度,进而影响泛化性能。实验结果表明,具有较大Fisher有效维度的电路在训练和测试之间存在更大的性能差距,而参数数量对泛化的预测能力较弱。该机制在多个任务中得到了验证,纠缠电路的泛化性能普遍低于相同参数数量的非纠缠电路。
🔬 方法详解
问题定义:本文旨在解决量子策略在泛化能力上的不确定性,现有方法主要依赖参数数量,未能有效解释泛化现象的根本原因。
核心思路:提出PAC-Bayesian框架,认为泛化能力受Fisher几何有效维度的影响,纠缠连接性作为复杂性的独立维度,影响电路的有效维度。
技术框架:研究通过控制实验,固定可训练旋转的数量,仅改变电路的纠缠程度,分析其对Fisher有效维度和泛化性能的影响。主要模块包括电路设计、有效维度计算和泛化性能评估。
关键创新:将纠缠作为影响量子策略泛化能力的独立因素,提出了一种新的评估电路性能的排名证书,能够有效区分相同参数数量的电路。
关键设计:在实验中,使用低方差决策模型,如单观察分类器和价值头,设计了多步策略学习的框架,确保了训练准确性和优化过程的有效性。
🖼️ 关键图片
📊 实验亮点
实验结果表明,具有较大Fisher有效维度的纠缠电路在训练和测试之间存在显著的性能差距,且在多个任务中,纠缠电路的泛化性能普遍低于相同参数数量的非纠缠电路,验证了理论框架的有效性。
🎯 应用场景
该研究为量子强化学习中的策略设计提供了新的视角,强调纠缠与泛化之间的权衡,具有重要的理论和实际应用价值。未来可在量子计算、量子博弈等领域推广应用,提升量子算法的泛化能力。
📄 摘要(原文)
Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning, yet it remains unclear when and why quantum policies generalize. We give a PAC-Bayesian account in which generalization is governed not by the raw number of circuit parameters, but by the effective dimension of the Fisher geometry induced by the circuit. This quantity is inflated by entanglement, making entangling connectivity an independent axis of complexity.In controlled experiments that fix the number of trainable rotations and vary only entanglement, we find that circuits with larger Fisher effective dimension exhibit larger train-test gaps, while parameter count is a weak predictor. The resulting bound acts primarily as a ranking certificate: it correctly orders circuits with identical parameter count, which parameter-counting bounds cannot do. We validate this mechanism across supervised classification, quantum contextual bandits, and value-function generalization, where entangled circuits consistently generalize worse than non-entangled circuits of equal parameter count, with gaps shrinking as sample size increases.Our strongest evidence comes from low-variance decision models, including single-observable classifiers, value heads, and one-step policies. In end-to-end multi-step policy learning, entanglement effects remain statistically significant but high return variance leaves the full ordering only partially resolved. Partial-correlation analysis shows that Fisher effective dimension screens off entangling pattern, and controls for training accuracy, readout, and optimizer rule out major optimization confounders. The effect also persists on an IBM Heron quantum processor under real noise. Overall, our results reframe quantum policy design around an entanglement--generalization trade-off rather than expressivity alone.