Co-design of LLM-based preference agents: participation may drive overtrust
作者: Michael J. Fell
分类: cs.CY, cs.AI, cs.HC
发布日期: 2026-07-23
备注: 43 pages (18 main text plus supplementary material), 5 figures
💡 一句话要点
探讨共同设计偏好代理的信任问题及其影响
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 偏好代理 共同设计 信任问题 人机对齐 定性研究 透明度
📋 核心要点
- 核心问题:现有的偏好代理设计方法可能导致人类与代理之间的系统性不对齐,影响信任度。
- 方法要点:通过与参与者共同设计偏好代理,探索参与过程对信任的影响,并提出透明度的重要性。
- 实验或效果:研究发现参与者对代理的信任度较高,但独立验证显示代理与人类反应存在显著差异。
📝 摘要(中文)
随着大型语言模型在研究和实际应用中越来越多地用于模拟人类偏好,关于验证、误表征和排除的担忧也随之增加。与被代表者共同设计代理是一种有前景的解决方案,但参与过程可能掩盖其表面上解决的问题。本文通过一项定性研究探讨了这一矛盾,12名参与者共同设计了家庭能源领域的个人偏好代理。尽管参与者普遍认为代理能够很好地代表他们,但独立验证显示人类与代理之间的对齐程度存在差异,代理的反应明显更为同质化、果断和抽象。本文认为,参与和过程透明度可能成为一种“过度信任引擎”,在促进信任的同时掩盖系统性不对齐的问题。
🔬 方法详解
问题定义:本文旨在解决大型语言模型在模拟人类偏好时可能导致的验证和信任问题。现有方法往往忽视了参与者与代理之间的系统性不对齐,可能导致过度信任。
核心思路:通过与参与者共同设计偏好代理,强调参与过程的透明度,以此来揭示潜在的信任问题和不对齐现象。
技术框架:研究采用定性方法,包括背景调查、共同设计访谈和验证调查三个主要阶段,确保参与者的声音被充分纳入设计过程。
关键创新:提出了“过度信任引擎”的概念,强调参与和透明度在偏好代理设计中的重要性,认为个体对齐是一个动态的过程,而非固定状态。
关键设计:在设计过程中,采用了多种访谈技术和调查问卷,以确保参与者的反馈能够有效反映在代理的设计中,同时关注代理反应的多样性和复杂性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,尽管参与者普遍认为代理能够准确代表他们的偏好,但独立验证表明,代理的反应在同质性和抽象性上明显高于人类样本。这一发现揭示了参与过程可能导致的信任偏差,强调了在设计过程中的透明度和验证的重要性。
🎯 应用场景
该研究的潜在应用领域包括智能家居、个性化推荐系统和人机交互等。通过共同设计偏好代理,可以提高用户的参与感和信任度,从而提升系统的实际应用效果和用户满意度。未来,该方法可能在更广泛的领域中推广,促进人机协作的有效性。
📄 摘要(原文)
Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an "overtrust engine" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.