cs.AI(2026-07-16)

📊 共 34 篇论文 | 🔗 1 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (25 🔗1) 支柱二:RL算法与架构 (RL & Architecture) (4) 支柱一:机器人控制 (Robot Control) (2) 支柱四:生成式动作 (Generative Motion) (1) 支柱八:物理动画 (Physics-based Animation) (1) 支柱三:空间感知与语义 (Perception & Semantics) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (25 篇)

#题目一句话要点标签🔗
1 VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence 提出VLT模型以解决工业时间序列多模态建模问题 large language model foundation model multimodal
2 Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy 基于科学可视化素养评估基准测试多模态大语言模型 large language model multimodal
3 CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models 提出CFM-Bench以解决CFM评估不统一的问题 foundation model multimodal
4 TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning 提出TopoAgent以解决多模态科学推理中的线性规划问题 large language model multimodal
5 MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization 提出MM-IssueLoc以解决多模态问题定位的评估挑战 multimodal
6 Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality 提出多语言提示下代码生成的基准研究以解决语言偏见问题 large language model
7 InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring 提出InCarEmo数据集以解决驾驶员情绪识别与状态监测问题 multimodal
8 Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction 提出记忆驱动的自我揭示模型以增强人机关系 multimodal
9 A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi 提出MiniCPM-V-4.6以实现高效的多模态推理 multimodal
10 Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation 提出Seer框架以加速多模态大语言模型推理效率 large language model multimodal
11 Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution 提出专用智能系统以优化BPMN图转化为可执行工作流 generalist agent large language model
12 RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning 提出RetroAgent以解决多步逆合成规划问题 large language model
13 SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration 提出SearchOS以解决信息检索代理协作中的任务跟踪问题 large language model
14 When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space 提出PRISM以解决语言模型在物理安全性中的风险问题 large language model
15 Plover: Steering GUI Agents through Plan-Centric Interaction 提出Plover以解决GUI自动化中的计划透明性问题 multimodal
16 Can We Trust Item Response Theory for AI Evaluation? 评估AI时提出IRT模型的局限性与改进建议 multimodal
17 Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation 提出基于GradientSHAP和隐式微分的优化推荐解释方法 large language model
18 Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience 利用LLMs从论文构建MaxSAT求解器的CoreForge经验 large language model
19 AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery 利用视觉语言模型推断建筑类型以辅助城市分析 chain-of-thought
20 SmartRAG: Native Graph-Based RAG for Mobile Device 提出SmartRAG以解决移动设备上大语言模型的计算挑战 large language model
21 LLM-Driven Approach to Modeling Tool Interoperability in Automotive Domain 提出LLM驱动的方法以解决汽车领域建模工具互操作性问题 large language model
22 Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments 提出多LLM代理的协作空间学习以优化网络效率 large language model
23 Towards an Intention Abstraction Layer for Autonomous Industrial Systems 提出意图抽象层以解决自主工业系统中的目标冲突问题 large language model
24 WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays 提出WrAFT以解决自动化写作评估的准确性与反馈问题 large language model
25 Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions 提出CEDI框架以解决多模态大语言模型评估的现实有效性问题 large language model

🔬 支柱二:RL算法与架构 (RL & Architecture) (4 篇)

#题目一句话要点标签🔗
26 Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment 提出Perception-RFT框架以解决多模态文档问答中的推理挑战 reinforcement learning multimodal visual grounding
27 Concept-Guided Spatial Regularization for World Models in Atari Pong 提出概念引导空间正则化以改善Atari Pong世界模型 reinforcement learning world model world models
28 Step-Level Preference Learning for Generative Agents in Social Simulations 提出步级偏好学习以提升社交模拟中的生成代理行为 preference learning direct preference optimization large language model
29 SMC-ES: Automated synthesis of formally verified control policies 提出SMC-ES以自动合成形式验证的控制策略 reinforcement learning deep reinforcement learning DRL

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
30 Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models 提出Action QFormer以优化视觉-语言-动作模型中的动作监督问题 sim-to-real vision-language-action VLA
31 Reachability-Aware Pretraining for Efficient Target-Oriented Path Exploration in Temporal Knowledge Graph Reasoning 提出RAPTOR以解决TKG推理中的路径探索效率问题 reachability-aware reinforcement learning

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
32 Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents 提出成本感知评估方法以优化安全代理的性能 penetration

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
33 Man, Machine, and Masterpiece: Artistic Ownership in the AI Era 提出ArtSplit以探讨AI时代艺术创作中的所有权问题 PULSE

🔬 支柱三:空间感知与语义 (Perception & Semantics) (1 篇)

#题目一句话要点标签🔗
34 Tactile: Giving Computer-Using Agents Hands and Feet 提出Tactile以解决计算机代理操作不可靠问题 affordance

⬅️ 返回 cs.AI 首页 · 🏠 返回主页