Creativity from Friction: Human-AI Interaction for Exploratory Structural Design

📄 arXiv: 2607.07521v1 📥 PDF

作者: Ricardo Maia Avelino, Rita Sevastjanova, Tom Van Mele, Philippe Block, Mennatallah El-Assady

分类: cs.HC, cs.AI

发布日期: 2026-07-08

备注: Accepted at ICML 2026, Workshop on Human-AI Co-Creativity


💡 一句话要点

提出人机交互系统以支持结构设计中的创意探索

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 人机交互 结构设计 生成式AI 视觉-语言模型 创意探索 设计空间 多模态系统

📋 核心要点

  1. 现有生成式AI系统往往无法满足结构设计等创意领域的需求,缺乏互动性和灵活性。
  2. 论文提出了一种基于视觉-语言模型的系统,旨在通过人机协作增强创意过程,支持设计师的迭代工作。
  3. 通过与专家的研究,表明该系统能够有效减少建模摩擦,提升设计师的创作体验和效率。

📝 摘要(中文)

本论文探讨了当前生成式AI在创意领域中的不足,特别是在结构设计和建筑领域。设计师需要能够外化和发展创意的互动系统,而不仅仅是提供最终答案。论文提出了一种基于视觉-语言模型的系统设计,旨在通过对话式、多模态的方式支持设计师的创作过程,减少重复建模的摩擦,同时保留反思设计的摩擦。通过与领域专家的研究,展示了该系统如何促进设计空间的探索。

🔬 方法详解

问题定义:论文要解决的问题是当前生成式AI系统在创意设计中的局限性,尤其是缺乏互动性和对设计师需求的响应。现有方法往往试图消除设计过程中的摩擦,但这在创意领域并不适用。

核心思路:论文的核心解决思路是设计一个能够支持人机协作的系统,使其能够通过对话式和多模态的方式与设计师互动,从而促进创意的生成和发展。这样的设计能够更好地适应设计师的迭代工作流程。

技术框架:整体架构包括用户输入、视觉-语言模型处理、设计空间探索和反馈机制等主要模块。系统通过分析用户的意图和需求,提供实时的设计建议和反馈。

关键创新:最重要的技术创新点在于将视觉-语言模型应用于结构设计的互动系统中,使得设计过程更加灵活和响应性强。这与传统的生成式AI方法有本质区别,后者通常只关注最终结果。

关键设计:在设计中,系统使用了特定的参数设置和损失函数,以确保模型能够有效理解和生成与结构设计相关的内容。此外,网络结构经过优化,以支持多模态输入和输出的处理。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,使用该系统的设计师在设计空间探索中表现出更高的效率和创造力。与传统方法相比,设计师在减少重复建模摩擦的同时,能够保持反思设计的深度,提升了整体设计质量和用户满意度。

🎯 应用场景

该研究的潜在应用领域包括建筑设计、工程设计和其他需要创意探索的领域。通过提供一个互动的AI系统,设计师能够更高效地探索设计空间,提升创意的实现能力,最终推动设计质量的提升。未来,该系统可能会在更多创意行业中得到应用,促进人机协作的进一步发展。

📄 摘要(原文)

AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design and architecture need interactive systems that help users externalise and develop ideas, explore alternatives, and refine partial solutions. The final product of such designs needs to comply with many constraints concerning, e.g., spatial configuration, mechanical behaviour, material quantities, and costs. These constraints create friction in the design process, which can stimulate novel and creative solutions. In this paper, we discuss the misalignment between current generative AI goals to remove friction and provide final solutions and the needs of creators, such as structural designers, who develop ideas through iterative work. We present the design dimensions of systems allowing for constrained human-AI co-creation that rely on vision-language models making structural exploration conversational, multimodal, and responsive to evolving human intent in ways that follow and augment the discipline's creative process. Through a pilot design interface based on these principles and a study with experts in the field, this paper shows how structural designers perceive interactive AI systems and how such systems can support design space exploration by reducing repetitive modelling friction while preserving reflective design friction.