A Comparative Analysis of Large Language Models on Clinical Questions for Autoimmune Diseases

Weiming Zhang, Jie Yu, Juntao Ma, Jiawei Feng,Linyu Geng,Yuxin Chen,Huayong Zhang,Mingzhe Ning

crossref（2024）

引用 0|浏览0

暂无评分

摘要

Background Artificial intelligence (AI) has made great strides. Our study evaluated the performance in delivering clinical questions related to autoimmune diseases (AIDs). Methods 46 AIDs-related questions were compiled and entered into ChatGPT 3.5, ChatGPT 4.0, and Gemini. The replies were collected and sent to laboratory specialists for scoring according to relevance, correctness, completeness, helpfulness, and safety. Scores for three chatbots in five quality dimensions and the scores of the replies to the questions under each quality dimension were analyzed. Results ChatGPT 4.0 showed superior performance than ChatGPT 3.5 and Gemini in all five quality dimensions. ChatGPT 4.0 outperformed ChatGPT 3.5 or Gemini on the relevance, completeness or helpfulness in answering about the prognosis, diagnosis, or the report interpretation of AIDs. ChatGPT 4.0’s replies were the longest, followed by ChatGPT 3.5, Gemini’s was the shortest. Conclusions Our findings highlight ChatGPT 4.0 is superior to delivering comprehensive and accurate responses to AIDs-related clinical questions.

查看译文

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要