此页面是自动翻译的,不保证翻译的准确性。请参阅 英文版 对于源文本。

Improving the Reliability of LLMs as Medical Assistants for the General Public (LAMP-1)

2026年9月1日 更新者:Ji Xunming,MD,PhD、Capital Medical University

Improving the Reliability of LLMs as Medical Assistants for the General Public: a Proof of Concept Simulation Trial

This study will evaluate whether three-minute six-dimensions education(3M-6D education) can improve the reliability of large language models as medical assistants for the general public. Participants will be randomly assigned to receive or not receive 3M-6D education and then use ChatGPT, Gemini, or non-AI information resources. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.

研究概览

详细说明

This randomized, controlled, proof-of-concept simulation trial will evaluate whether three-minute six-dimensions education (3M-6D education) can improve the reliability of large language models as medical assistants for the general public.

Eligible participants will be randomly assigned in a 1:1:1:1:1 ratio to one of five study groups: the 3M-6D education GPT group, the GPT group, the 3M-6D education Gemini group, the Gemini group, or the control group. Participants in the 3M-6D education GPT and 3M-6D education Gemini groups will receive approximately three minutes of education before using ChatGPT or Gemini.Each participant will be randomly assigned one of 10 standardized clinical scenarios and complete a simulated counseling task in unrestricted natural language within approximately 10 minutes. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.

研究类型

介入性

注册 (实际的)

527

阶段

  • 不适用

联系人和位置

本节提供了进行研究的人员的详细联系信息,以及有关进行该研究的地点的信息。

学习地点

    • Beijing Municipality
      • Beijing、Beijing Municipality、中国
        • Beijing Ctiy

参与标准

研究人员寻找符合特定描述的人,称为资格标准。这些标准的一些例子是一个人的一般健康状况或先前的治疗。

资格标准

适合学习的年龄

  • 成人
  • 年长者

接受健康志愿者

是的

描述

Inclusion Criteria:

  1. Age 18 years or greater, male or female;
  2. Completed primary school or higher education;
  3. Able to use a smartphone or computer to complete online interaction;
  4. No history of acute ischemic stroke, systemic lupus erythematosus, gastric ulcer, pneumonia, acute cardiac infarction, urinary tract infection, uterine fibroids, diabetes, osteoarthritis, or migraine.
  5. Able to understand and comply with study procedures and to provide written informed consent.

Exclusion Criteria:

  1. Currently or previously employed as a healthcare worker;
  2. Previously received systematic medical training;
  3. Currently involved in concurrent research that may interfere with the results of the present trial;
  4. The investigator considered that the participant had other conditions that might affect compliance or preclude participation.

学习计划

本节提供研究计划的详细信息,包括研究的设计方式和研究的衡量标准。

研究是如何设计的?

设计细节

  • 主要用途:卫生服务研究
  • 分配:随机化
  • 介入模型:并行分配
  • 屏蔽:单身的

武器和干预

参与者组/臂
干预/治疗
实验性的:3M-6D education GPT Group
Participants will first be trained in 3M-6D education, then use ChatGPT to complete a consultation task in unrestricted natural language in approximately 10 minutes.
Participants use ChatGPT to complete a standardized simulated clinical scenarios in unrestricted natural language.

3M-6D education is designed based on Cognitive Load Theory to reduce the cognitive burden on patients during medical interactions with AI and to improve the clarity and completeness of symptom reporting.

Guided by cognitive load theory and the natural process physicians use to take medical histories, the investigators identified candidate information dimensions and developed a structured expression framework with six dimensions for public health queries through a Delphi expert consensus process. Participants were instructed to use the framework to describe their symptoms across these six dimensions; this process can typically be completed within three minutes, so the investigators call this approach three minutes six dimensions education (3M-6D education).

实验性的:3M-6D education Gemini Group
Participants will first be trained in 3M-6D education, then use Gemini to complete a consultation task in unrestricted natural language in approximately 10 minutes.
Participants use Gemini to complete a standardized simulated clinical scenarios in unrestricted natural language.

3M-6D education is designed based on Cognitive Load Theory to reduce the cognitive burden on patients during medical interactions with AI and to improve the clarity and completeness of symptom reporting.

Guided by cognitive load theory and the natural process physicians use to take medical histories, the investigators identified candidate information dimensions and developed a structured expression framework with six dimensions for public health queries through a Delphi expert consensus process. Participants were instructed to use the framework to describe their symptoms across these six dimensions; this process can typically be completed within three minutes, so the investigators call this approach three minutes six dimensions education (3M-6D education).

有源比较器:GPT Group
Participants will use ChatGPT to complete a consultation task in unrestricted natural language in approximately 10 minutes.
Participants use ChatGPT to complete a standardized simulated clinical scenarios in unrestricted natural language.
有源比较器:Gemini Group
Participants will use Gemini to complete a consultation task in unrestricted natural language in approximately 10 minutes.
Participants use Gemini to complete a standardized simulated clinical scenarios in unrestricted natural language.
无干预:Control group
Participants will use non-AI tools such as internet searches and medical websites to complete a consultation task in unrestricted natural language in approximately 10 minutes.

研究衡量的是什么?

主要结果指标

结果测量
措施说明
大体时间
Relevant conditions identification of the 3M-6D education GPT group compared with the GPT group
大体时间:1 hour.
Relevant conditions identification is defined as the proportion of participants whose final response includes the expert-defined final diagnosis or a relevant differential diagnosis.
1 hour.
Disposition concordance of the 3M-6D education GPT group compared with the GPT group
大体时间:1 hour.
Disposition concordance is defined as the proportion of participants whose final care recommendation matches the expert-defined level. The five levels are self-care, routine outpatient care, urgent outpatient care, emergency department visit, and emergency medical services.
1 hour.
Relevant conditions identification of the 3M-6D education Gemini group compared with the Gemini group
大体时间:1 hour.
1 hour.
Disposition concordance of the 3M-6D education Gemini group compared with the Gemini group
大体时间:1 hour.
1 hour.

次要结果测量

结果测量
措施说明
大体时间
Relevant conditions identification of the 3M-6D education GPT group compared with the control group
大体时间:1 hour.
1 hour.
Relevant conditions identification of the 3M-6D education Gemini group compared with the control group
大体时间:1 hour.
1 hour.
Disposition concordance of the 3M-6D education GPT group compared with the control group
大体时间:1 hour.
1 hour.
Disposition concordance of the 3M-6D education Gemini group compared with the control group
大体时间:1 hour.
1 hour.
Red-flag identification in the 3M-6D education GPT group compared with the GPT group
大体时间:1 hour.
Red-flag identification is defined as the proportion of participants whose final response includes the key warning signs that experts defined for the assigned scenario.
1 hour.
Red-flag identification in the 3M-6D education GPT group compared with the control group
大体时间:1 hour.
1 hour.
Red-flag identification in the 3M-6D education Gemini group compared with the Gemini group
大体时间:1 hour.
1 hour.
Red-flag identification in the 3M-6D education Gemini group compared with the control group
大体时间:1 hour.
1 hour.
NASA Task Load Index score of the 3M-6D education GPT group compared with the GPT group
大体时间:1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
1 hour.
NASA Task Load Index score of the 3M-6D education GPT group compared with the control group
大体时间:1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
1 hour.
NASA Task Load Index score of the 3M-6D education Gemini group compared with the Gemini group
大体时间:1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
1 hour.
NASA Task Load Index score of the 3M-6D education Gemini group compared with the control group
大体时间:1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
1 hour.
Relevant conditions identification of the 3M-6D education GPT group compared with the 3M-6D education Gemini group
大体时间:1 hour.
1 hour.
Disposition concordance of the 3M-6D education GPT group compared with the 3M-6D education Gemini group
大体时间:1 hour.
1 hour.
Red-flag identification in the 3M-6D education GPT group compared with the 3M-6D education Gemini group
大体时间:1 hour.
1 hour.
NASA Task Load Index score of the 3M-6D education GPT group compared with the 3M-6D education Gemini group
大体时间:1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
1 hour.

其他结果措施

结果测量
措施说明
大体时间
Failure to identify red flags in the 3M-6D education GPT group compared with the GPT group
大体时间:1 hour.
Failure to identify red flags is defined as the proportion of participants whose final response does not include the expert-defined red-flag symptoms or warning signs for the assigned standardized simulated clinical scenario.
1 hour.
Failure to identify red flags in the 3M-6D education GPT group compared with the control group
大体时间:1 hour.
1 hour.
Failure to identify red flags in the 3M-6D education Gemini group compared with the Gemini group
大体时间:1 hour.
1 hour.
Failure to identify red flags in the 3M-6D education Gemini group compared with the control group
大体时间:1 hour.
1 hour.
Underestimation of disposition in the 3M-6D education GPT group compared with the GPT group
大体时间:1 hour.
Underestimation of disposition is defined as the proportion of participants whose final care recommendation is lower than the expert-defined disposition level for the assigned standardized simulated clinical scenario.
1 hour.
Underestimation of disposition in the 3M-6D education GPT group compared with the control group
大体时间:1 hour.
1 hour.
Underestimation of disposition in the 3M-6D education Gemini group compared with the Gemini group
大体时间:1 hour.
1 hour.
Underestimation of disposition in the 3M-6D education Gemini group compared with the control group
大体时间:1 hour.
1 hour.

合作者和调查者

在这里您可以找到参与这项研究的人员和组织。

研究记录日期

这些日期跟踪向 ClinicalTrials.gov 提交研究记录和摘要结果的进度。研究记录和报告的结果由国家医学图书馆 (NLM) 审查,以确保它们在发布到公共网站之前符合特定的质量控制标准。

研究主要日期

学习开始 (实际的)

2026年7月3日

初级完成 (实际的)

2026年8月26日

研究完成 (实际的)

2026年8月26日

研究注册日期

首次提交

2026年6月11日

首先提交符合 QC 标准的

2026年6月11日

首次发布 (实际的)

2026年6月16日

研究记录更新

最后更新发布 (实际的)

2026年9月3日

上次提交的符合 QC 标准的更新

2026年9月1日

最后验证

2026年9月1日

更多信息

与本研究相关的术语

计划个人参与者数据 (IPD)

计划共享个人参与者数据 (IPD)?

未定

药物和器械信息、研究文件

研究美国 FDA 监管的药品

研究美国 FDA 监管的设备产品

此信息直接从 clinicaltrials.gov 网站检索,没有任何更改。如果您有任何更改、删除或更新研究详细信息的请求,请联系 register@clinicaltrials.gov. clinicaltrials.gov 上实施更改,我们的网站上也会自动更新.

订阅