| 1 |
Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations |
提出推理一致性扫描框架以审计AI安全评估中的思维链有效性 |
chain-of-thought |
|
|
| 2 |
Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies |
综述LLM在网络安全中的双重应用与防御策略 |
large language model |
|
|
| 3 |
Physics-Audited Agentic Discovery in Scientific Machine Learning |
提出物理审核的代理科学机器学习以解决模型验证问题 |
large language model |
|
|
| 4 |
Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass |
提出Jailbreak以解决数据库访问瓶颈问题 |
large language model |
|
|
| 5 |
Creativity from Friction: Human-AI Interaction for Exploratory Structural Design |
提出人机交互系统以支持结构设计中的创意探索 |
multimodal |
|
|
| 6 |
From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents |
提出EvoSOP框架以优化LLM代理的工具使用效率 |
large language model |
|
|
| 7 |
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks |
提出ImagingBench基准以评估代理AI在计算成像任务中的表现 |
multimodal |
|
|
| 8 |
Learning social norms enhances compatibility in dynamic human-AI coordination |
通过学习社会规范提升动态人机协调的兼容性 |
large language model |
|
|
| 9 |
MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations |
提出MADB数据集以解决音乐美学评估的挑战 |
multimodal |
✅ |
|
| 10 |
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI |
提出Harness设计以优化企业智能AI的代币经济 |
foundation model |
|
|