cs.AI(2026-07-20)

📊 共 79 篇论文

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (51) 支柱二:RL算法与架构 (RL & Architecture) (22) 支柱一:机器人控制 (Robot Control) (3) 支柱六:视频提取与匹配 (Video Extraction) (1) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱四:生成式动作 (Generative Motion) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (51 篇)

#题目一句话要点标签🔗
1 FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment 提出FAIR_XAI以解决多模态模型公平性问题 foundation model multimodal
2 Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length 提出LaRT模型以提升大语言模型评估的准确性与效率 large language model chain-of-thought
3 Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization 提出SigMap以解决跨场景无线定位问题 foundation model multimodal
4 MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion 提出MGDT以解决多模态知识图谱补全问题 large language model multimodal
5 Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts 提出神经符号AI以解决LEED合规性验证问题 multimodal chain-of-thought
6 Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal 提出链式思维干扰以解决拒绝控制问题 chain-of-thought
7 S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation 提出S1-Omni以解决科学理解与预测的多模态建模问题 multimodal
8 SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery 提出SciForge以解决科学研究中的多模态协作问题 multimodal
9 The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning 提出TRACE框架以增强多模态推理的证据基础 multimodal
10 Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models 提出低可恢复性隐写术以应对大语言模型中的隐写攻击 large language model
11 RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources 提出RESOURCE2SKILL框架以从多模态资源中提炼可执行代理技能 multimodal
12 Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models 提出隐私保护的边缘云协作推理框架以解决大语言模型推理问题 large language model
13 Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction 提出记忆驱动的自我披露模型以增强人机关系 multimodal
14 DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings 提出DrawingVQA以评估多模态大语言模型在施工图上的推理能力 large language model multimodal
15 SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents 提出SciVisAgentBench以评估科学数据分析与可视化代理的能力 large language model multimodal
16 Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents 提出反思性代理模型以提升信息提取任务的可控性 large language model
17 NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning 提出NeurOWL以解决不完整OWL本体推理问题 large language model
18 Knowledge-Centric Agents for Workflow Generation 提出知识中心代理以解决工作流生成问题 large language model
19 Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation 提出管理者强制基准以评估AI间的管理动态 chain-of-thought
20 Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design 提出进化算法引导的LLMs以优化物理信息神经网络设计 large language model
21 LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization 提出基于大型语言模型的自主AI以优化5G/6G网络管理 large language model
22 Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities 评估开放权重LLM生成自动驾驶车辆漏洞的结构化威胁信息 large language model
23 Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy 通过线性探测揭示大语言模型的认知复杂性 large language model
24 Towards a General Intelligence and Interface for Wearable Health Data 提出可穿戴健康数据的基础模型以解决个性化健康洞察问题 foundation model
25 Agents-K1: Towards Agent-native Knowledge Orchestration 提出Agents-K1以解决科学知识编排不足问题 multimodal
26 Can We Trust Item Response Theory for AI Evaluation? 探讨IRT在AI评估中的可靠性挑战与解决方案 multimodal
27 Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios 提出CLI-Tool-Bench以评估LLM驱动的从零到一软件生成能力 large language model
28 Energy-based Transport for Amortized Bayesian Inference 提出基于能量的传输方法以解决非线性逆问题的贝叶斯推断 multimodal
29 Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations 评估vLLM配置对能耗、性能与准确性的影响 large language model
30 Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs 提出Code-MUE以解决代码LLMs不确定性评估问题 large language model
31 PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model 提出PGN以解决视觉语言导航中的多模态对齐问题 VLN large language model foundation model
32 ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding 提出ST-Veto以提升扩散多模态大语言模型的推理能力 large language model multimodal visual grounding
33 Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation 提出Pailitao-MMSearch以解决电商多模态搜索问题 large language model foundation model multimodal
34 Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding 提出ChoroplethMap-Bench以提升机器空间理解能力 foundation model
35 Human Grounded Evaluation of Large Language Models for Optical Network Automation 提出HuGLEN以解决光网络自动化中的LLM评估问题 large language model
36 Stress Testing Concept Erasure with Large Language Model Agents 提出STACE框架以解决概念消除评估的挑战 large language model
37 WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement 提出WuYu-EnvLE-Bench以评估环境执法中的大型语言模型 large language model
38 SGA: Plug&Play Geometric Verification for Educational Video Synthesis 提出SGA模块以解决教育视频合成中的几何验证问题 large language model
39 Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows 提出FluxBench以评估AI代理在EDA工作流中的表现 foundation model
40 SALT: Salience-Aware Lexical Trie for Long-Context Compression 提出SALT以解决长上下文压缩中的主题覆盖问题 large language model
41 Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering 提出SOPHIA以解决大型语言模型自循环推理问题 large language model
42 AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models 提出AdaHome以解决智能家居助手的效率与隐私问题 large language model
43 Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation 提出自然语言访问领域特定元数据框架以解决查询生成问题 large language model
44 OntoExtend: A Framework for Requirement-driven and Scalable Ontology Extension with LLMs 提出OntoExtend框架以解决本体扩展中的需求驱动问题 large language model
45 Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking 提出基于LLM的代理模型以提升ABM的可靠性与性能 large language model
46 Persona-as-Configuration: Generative Stakeholder Reporting for Agricultural Floods 提出Persona-as-Configuration以解决农业洪水报告生成问题 large language model
47 Autonomous Discovery of Wireless Communications Algorithms 提出AITE框架以自主设计无线通信算法 large language model
48 LaT: LLM-as-Trainer for Multi-Task Vehicle Routing Solvers 提出LLM-as-Trainer以解决多任务车辆路径规划问题 large language model
49 FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models 提出FlowBlock以解决自校正扩散语言模型的解码效率问题 large language model
50 Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents 提出VRR-Stop框架以解决LLM代理中的噪声验证-修复循环问题 large language model
51 Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory 提出预算依赖的操作选择机制以优化语言代理的记忆管理 large language model

🔬 支柱二:RL算法与架构 (RL & Architecture) (22 篇)

#题目一句话要点标签🔗
52 JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models 提出JoyNexus以解决多租户VLA模型后训练效率问题 reinforcement learning vision-language-action VLA
53 Length Penalties Make Chain-of-Thought Less Monitorable 提出长度惩罚以提高链式推理的可监控性 reinforcement learning chain-of-thought
54 DSWorld: A Data Science World Model for Efficient Autonomous Agents 提出DSWorld以提升自主数据科学代理的效率 reinforcement learning world model world models
55 AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis 提出AuEmoChat以解决对话语音合成中的情感真实度问题 flow matching motion representation multimodal
56 Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3? 提出多种编码代理以解决ARC-AGI-3性能归因问题 world model world models
57 SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction 提出SeerGuard框架以解决移动GUI代理的安全风险问题 world model world models
58 AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning 提出AV-JEPA以解决音视频自监督学习的对齐问题 JEPA multimodal
59 ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning 提出ToolVerse以解决大规模环境下长时间任务的挑战 reinforcement learning world model world models
60 Human-Aligned Procedural Level Generation Reinforcement Learning via Text-Level-Sketch Shared Representation 提出VIPCGRL以解决人类中心的程序内容生成问题 reinforcement learning deep reinforcement learning contrastive learning
61 RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs 提出RL-Struct框架以解决LLM生成与结构化需求之间的差距 reinforcement learning PPO
62 From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems 通过Prolog专家系统实现可解释的深度强化学习 reinforcement learning deep reinforcement learning
63 UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation 提出UCOB框架以提升强化学习中的技能利用与演化 reinforcement learning distillation
64 RAD: Retrieval High-quality Demonstrations to Enhance Decision-making 提出RAD以解决离线强化学习中的泛化问题 reinforcement learning offline RL offline reinforcement learning
65 ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning 提出ACPO以解决多智能体强化学习中的策略优化问题 reinforcement learning
66 Reinforcement Learning: From Algorithms To Foundation Models 提出多智能体强化学习与基础模型结合的方法以解决复杂决策问题 reinforcement learning world model world models
67 OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment 提出OrientSAM以解决多模态空间推理中的相机中心快捷方式问题 curriculum learning large language model multimodal
68 Mobile Network Control with a World Model 提出基于世界模型的移动网络控制方法以提升能效 reinforcement learning world model world models
69 SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning 提出SAGE以解决长时间规划中的候选质量问题 world model world models
70 Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents? 提出反馈增强自蒸馏以解决检索交错搜索代理的性能问题 distillation privileged information large language model
71 Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective 提出基于加权模型平均的异构大语言模型合并方法 distillation large language model instruction following
72 PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning 提出PAMD以解决视觉强化学习中的距离度量问题 reinforcement learning
73 Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution 提出跨模态注意力架构以解决否定检测问题 representation learning multimodal

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
74 HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization 提出HiLSVA以解决科学可视化中的人机协作问题 manipulation large language model
75 Perceived AGI: Believability as Dimensional Completeness, Not Capability 提出维度完整性以提升人工对话体的可信度 manipulation large language model
76 The AI Fiction Paradox 提出AI小说悖论以解决长篇小说生成问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
77 Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes 提出MAR-12框架以解决有害幽默检测问题 HuMoR multimodal

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
78 ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory 提出ABot-AgentOS以解决长时间嵌入式智能体的推理与记忆问题 cross-embodiment VLA

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
79 Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents 提出成本感知评估方法以优化安全代理的性能 penetration

⬅️ 返回 cs.AI 首页 · 🏠 返回主页