Dimensionality in Satisfaction Ratings

📄 arXiv: 2607.11026v1 📥 PDF

作者: Andrew Hong, Jason Potteiger

分类: cs.CL

发布日期: 2026-07-13

备注: 25 pages, 7 figures, 6 tables


💡 一句话要点

利用大语言模型分解客户满意度以提升体验分析

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 客户满意度 大语言模型 维度分析 用户体验 数据分析 机器学习 自然语言处理

📋 核心要点

  1. 现有方法在客户满意度评估中缺乏细致的维度分析,导致对客户体验的理解不够全面。
  2. 本研究提出利用大语言模型对客户支持对话进行注释,分解满意度为多个维度,以便更好地理解客户体验。
  3. 实验结果表明,分解后的满意度维度与客户自我报告的满意度高度相关,且提供了更深入的客户体验洞察。

📝 摘要(中文)

本研究使用大型语言模型(GPT-4.1)对全球消费品公司的约9000个支持对话文本进行注释,将客户关怀满意度分解为多个维度(整体、代理、结果、产品和客户努力),并验证了模型注释与客户自我报告满意度之间的关系。结果显示,四个维度与自我报告的满意度高度相关,而产品满意度的相关性较弱。通过排除严重分歧的会话,整体相关性显著提高,表明分解满意度的价值在于识别客户体验的细微驱动因素。

🔬 方法详解

问题定义:本研究旨在解决客户满意度评估中缺乏细致维度分析的问题。现有方法往往只关注整体满意度,忽视了影响客户体验的多种因素。

核心思路:通过使用大型语言模型对客户支持对话进行注释,将满意度分解为多个维度,以识别更细致的客户体验驱动因素。这种设计能够提供更全面的满意度分析。

技术框架:整体流程包括数据收集、模型注释、维度分解和相关性验证。主要模块包括对话文本的预处理、GPT-4.1模型的应用以及满意度维度的统计分析。

关键创新:本研究的创新点在于首次将大型语言模型应用于客户满意度的多维度分析,提供了比传统方法更深入的客户体验洞察。与现有方法相比,强调了维度分解的价值在于归因和覆盖,而非单纯的预测。

关键设计:在模型注释过程中,设置了多个参数以优化语言模型的输出,使用了适当的损失函数来提高注释的准确性。模型的训练和验证过程中,特别关注了与客户自我报告满意度的相关性。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,四个维度的满意度与客户自我报告的满意度相关性高达0.65,而在排除严重分歧后,整体相关性提升至0.914。这表明分解满意度的分析方法能够显著提高对客户体验的理解。

🎯 应用场景

该研究的潜在应用领域包括客户服务、市场研究和用户体验设计等。通过更细致的满意度分析,企业能够更好地识别客户需求和痛点,从而优化服务流程和产品设计,提升客户满意度和忠诚度。

📄 摘要(原文)

We used a large language model (GPT-4.1) to annotate the text of about 9,000 support conversations at a global consumer-goods firm, decomposing customer-care satisfaction into component axes (overall, agent, outcome, product, and customer effort), and validated the LLM annotations against the satisfaction ratings customers gave themselves. Four of five axes track self-reported satisfaction closely (overall, agent, and outcome near an unadjusted 0.65; effort -0.54), while product satisfaction is weak against the available proxy. The unadjusted correlation also understates the alignment: the disagreements concentrate in a small, readable tail of divergent sessions rather than in general drift, and the overall correlation rises to 0.811 when only the severe divergences are excluded and to 0.914 when the full divergent tail is excluded. The axes are also highly collinear, and adding them to the overall score does not improve prediction of the customer's rating, the decomposition's value is not incremental prediction but attribution and coverage. And, with greater coverage the picture of the data changes. Read on every contact rather than the few that return a survey, satisfaction is markedly lower than the survey reports (a full-census 2.91 against the surveyed 3.62 on a five-point scale). The promise of decomposed satisfaction as a methodology is the ability to identify more nuanced drivers of customer experience in conversational data.