Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development

📄 arXiv: 2607.18696v1 📥 PDF

作者: Yinan Wang

分类: cs.AI

发布日期: 2026-07-21

备注: Working paper. Includes public no-label benchmark cases and dry-lab evaluation artifacts. No wet-lab, patient-level, clinical, regulatory, or investment validation is claimed


💡 一句话要点

提出公司世界模型以优化AI驱动药物开发

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 生物技术 AI驱动 药物开发 公司世界模型 决策优化 资产价值 动态更新

📋 核心要点

  1. 现有的AI原生生物技术公司往往依赖于传统的组织结构,导致效率低下和创新不足。
  2. 本文提出的公司世界模型通过动态的资产-价值状态表示,优化了决策过程,超越了传统部门划分。
  3. 实验结果表明,价值转换架构在自动评分中表现优异,且在盲评中获得了更高的认可度。

📝 摘要(中文)

AI原生生物技术公司通常通过模仿人类生物技术组织结构来设计代理角色。本文提出了一种不同的抽象概念:公司世界模型,定义为一个持久的资产-价值状态表示,包含转移模型、明确的价值函数、规划和更新,涵盖科学、监管、商业开发、财务和执行约束。我们引入了一个干实验基准,用于测试AI代理组织是否应模仿部门或围绕这种世界模型运作。基准包含45个回顾性公共信息决策案例,具有严格的时间截止、隐藏结果、共同模式、自动评分和盲评。研究比较了人类组织模仿、增强人类组织模仿、AI原生资产中心和AI原生价值转换架构。结果显示,价值转换架构在自动价值转换评分中表现最佳,并受到价值特定盲评审的强烈偏好。

🔬 方法详解

问题定义:本文旨在解决AI原生生物技术公司在组织结构设计上的不足,现有方法往往依赖于静态的人类组织图,导致决策效率低下和创新能力不足。

核心思路:提出公司世界模型作为一种动态的资产-价值状态表示,强调在科学、监管和商业等多重约束下的决策优化,旨在提升AI代理的决策能力。

技术框架:整体架构包括资产-价值状态表示、转移模型、价值函数、规划模块和更新机制,涵盖了决策的各个方面。

关键创新:最重要的技术创新在于提出了价值转换架构,它通过实时更新的资产价值记录,超越了传统的组织结构,提供了一种更为灵活和高效的决策支持系统。

关键设计:在设计中,采用了Deal、Approval、Revenue和Investment Arbiter循环来动态更新资产价值,确保决策过程的实时性和准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,价值转换架构在自动价值转换评分中取得了最高分,且在价值特定的盲评中被强烈偏好,表明其在决策优化方面的显著优势。与传统方法相比,表现出更高的灵活性和适应性。

🎯 应用场景

该研究的潜在应用领域包括药物开发、临床试验设计和生物技术公司的战略规划。通过优化决策过程,能够提高药物开发的效率和成功率,具有重要的实际价值和未来影响。

📄 摘要(原文)

AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles. We argue for a different abstraction: a Company World Model, defined as a persistent asset-to-value state representation with transition models, explicit value functions, planning, and updating across scientific, regulatory, BD, commercial, financial, and execution constraints. We introduce a dry-lab benchmark for testing whether AI-agent organizations should mimic departments or operate around such a world model. The benchmark contains 45 retrospective public-information decision cases with strict time cutoffs, hidden outcomes, common schemas, automatic scoring, and blinded pairwise judging. We compare human-org-mimic, stronger human-org-mimic-plus, AI-native asset-centric, and AI-native value-conversion architectures. The value-conversion architecture is a prompt-level approximation of a Company World Model: a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops. Under a success function defined by external BD, regulatory approval and launch, and revenue discipline, it achieved the highest automatic value-conversion score and was strongly preferred over the original baselines by value-specific blinded judges. Stress tests narrowed the claim: a stronger human baseline remained competitive, and a neutral judge did not show robust value-conversion dominance. Codex-only mechanistic ablations suggest that Revenue Room, Deal Room, and Approval Room carry useful work under the target objective. The central finding is objective-sensitive: departments may remain useful governance views, but the core AI-native operating primitive should be a shared, predictive asset-to-value state rather than a static human org chart. The study is dry-lab only and does not establish real-world drug success, clinical benefit, or revenue prediction accuracy.