| 1 |
JudgeLM: Fine-tuned Large Language Models are Scalable Judges |
提出JudgeLM以高效评估大语言模型在开放场景中的表现 |
large language model multimodal |
✅ |
|
| 2 |
Incorporating Probing Signals into Multimodal Machine Translation via Visual Question-Answering Pairs |
提出通过视觉问答对增强多模态机器翻译的交互性 |
large language model multimodal |
✅ |
|
| 3 |
ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-Thought |
提出ACT-SQL以提升文本到SQL的推理能力 |
large language model chain-of-thought |
|
|
| 4 |
M2C: Towards Automatic Multimodal Manga Complement |
提出多模态漫画补全方法以解决漫画理解问题 |
large language model multimodal |
|
|
| 5 |
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages |
评估GPT-4在系统评价过程中的应用潜力 |
large language model |
|
|
| 6 |
Evaluation of large language models using an Indian language LGBTI+ lexicon |
提出基于印度语言LGBTI+词典的LLM评估方法 |
large language model |
|
|
| 7 |
An Open Source Data Contamination Report for Large Language Models |
提出开放源代码数据污染报告以解决大语言模型评估问题 |
large language model |
|
|
| 8 |
The impact of responding to patient messages with large language model assistance |
利用大型语言模型辅助响应患者信息以减轻临床负担 |
large language model |
|
|
| 9 |
InstOptima: Evolutionary Multi-objective Instruction Optimization via Large Language Model-based Instruction Operators |
提出InstOptima以解决指令生成效率低下问题 |
large language model |
|
|
| 10 |
Proving Test Set Contamination in Black Box Language Models |
提出无需访问预训练数据的语言模型测试集污染证明方法 |
large language model |
|
|
| 11 |
"You Are An Expert Linguistic Annotator": Limits of LLMs as Analyzers of Abstract Meaning Representation |
评估大型语言模型在抽象意义表示分析中的局限性 |
large language model |
|
|
| 12 |
Outlier Dimensions Encode Task-Specific Knowledge |
研究表明异常维度编码任务特定知识 |
large language model |
|
|
| 13 |
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset |
利用大型语言模型自动评分短答案阅读理解题 |
large language model |
|
|
| 14 |
Symbolic Planning and Code Generation for Grounded Dialogue |
提出模块化对话系统以解决任务导向对话中的局限性 |
large language model |
|
|
| 15 |
A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications |
提出自动化框架以测量生成式AI应用中的责任AI危害 |
large language model |
|
|
| 16 |
An Ensemble Method Based on the Combination of Transformers with Convolutional Neural Networks to Detect Artificially Generated Text |
提出基于变换器与卷积神经网络的集成方法以检测人工生成文本 |
large language model |
|
|
| 17 |
X-SNS: Cross-Lingual Transfer Prediction through Sub-Network Similarity |
提出基于子网络相似性的方法以提升跨语言迁移预测能力 |
foundation model |
|
|
| 18 |
Techniques for supercharging academic writing with generative AI |
提出人机协作框架以提升学术写作质量与效率 |
large language model |
|
|
| 19 |
FLEEK: Factual Error Detection and Correction with Evidence Retrieved from External Knowledge |
提出FLEEK以解决文本信息中的事实错误检测与修正问题 |
large language model |
|
|