| 1 |
ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI |
提出ORCH以优化多智能体系统的组织与协调 |
embodied AI large language model |
|
|
| 2 |
SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control |
提出SIRF模型以解决工业内容风险控制问题 |
foundation model chain-of-thought |
|
|
| 3 |
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents |
提出Mr.LHDR基准以评估长时段多模态深度研究代理的能力 |
multimodal |
|
|
| 4 |
When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting |
提出信息论指标基准以评估文本在多模态时间序列预测中的贡献 |
multimodal |
|
|
| 5 |
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents |
提出Sci-MMR以解决多步骤证据基础科学推理问题 |
multimodal |
|
|
| 6 |
Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training |
提出多代理大语言模型系统以提升临床访谈训练效果 |
large language model |
|
|
| 7 |
A Mathematical Theory of Pragmatic Information |
提出实用信息理论以统一通信、控制与决策问题 |
embodied AI |
|
|
| 8 |
Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models |
提出对话式XAI助手以解决能源消费模型解释问题 |
large language model |
|
|
| 9 |
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization |
提出COBRA-Skills以解决技能优化中的高成本问题 |
large language model |
|
|
| 10 |
Characterizing Job Power Elasticity for Power-Flexible AI Training |
提出作业功率弹性指标以优化AI训练的电力使用 |
large language model |
|
|
| 11 |
ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps |
提出ActMap以解决大语言模型的不确定性量化问题 |
large language model |
|
|
| 12 |
Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning |
提出稳定性意识的测试时适应方法以提升LLM推理能力 |
large language model |
|
|
| 13 |
AI Soccer Analyst: Stage-Aware and Verifiable Human-AI Collaboration for Soccer Data Analysis |
提出AI足球分析师以解决数据分析中的人机协作问题 |
large language model |
|
|
| 14 |
(Whose defaults?) Is artificial intelligence reorienting archaeological methods? |
评估大型语言模型对考古方法多样性的影响 |
large language model |
|
|
| 15 |
SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics |
提出SemVerBench以评估LLM对版本约束解析语义的理解 |
large language model |
|
|
| 16 |
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation |
提出Benchmark Radar以解决AI基准评估检索问题 |
large language model |
|
|
| 17 |
Demystifying the Privacy-Utility Trade-off in LLM Interactions |
提出意图驱动的本地保护框架以优化隐私与效用的权衡 |
large language model |
|
|