Improving the Reliability of LLMs as Medical Assistants for the General Public (LAMP-1)
Improving the Reliability of LLMs as Medical Assistants for the General Public: a Proof of Concept Simulation Trial
Aperçu de l'étude
Statut
Statut
Les conditions
Les conditions
Intervention / Traitement
Intervention / Traitement
Description détaillée
This randomized, controlled, proof-of-concept simulation trial will evaluate whether three-minute six-dimensions education (3M-6D education) can improve the reliability of large language models as medical assistants for the general public.
Eligible participants will be randomly assigned in a 1:1:1:1:1 ratio to one of five study groups: the 3M-6D education GPT group, the GPT group, the 3M-6D education Gemini group, the Gemini group, or the control group. Participants in the 3M-6D education GPT and 3M-6D education Gemini groups will receive approximately three minutes of education before using ChatGPT or Gemini.Each participant will be randomly assigned one of 10 standardized clinical scenarios and complete a simulated counseling task in unrestricted natural language within approximately 10 minutes. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.
Type d'étude
Type d'étude
Inscription (Réel)
Inscription
Phase
Phase
- N'est pas applicable
Contacts et emplacements
Coordonnées de l'étude
Coordonnées de l'étude
- Nom: Xunming Ji
- Numéro de téléphone: 01083198962
- E-mail: jixm@ccmu.edu.cn
Sauvegarde des contacts de l'étude
- Nom: Chuanjie Wu
- Numéro de téléphone: 01083199439
- E-mail: wuchuanjie@ccmu.edu.cn
Lieux d'étude
-
-
Beijing Municipality
-
Beijing, Beijing Municipality, Chine
- Beijing Ctiy
-
-
Critères de participation
Critère d'éligibilité
Critère d'éligibilité
Âges éligibles pour étudier
- Adulte
- Adulte plus âgé
Accepte les volontaires sains
La description
Inclusion Criteria:
- Age 18 years or greater, male or female;
- Completed primary school or higher education;
- Able to use a smartphone or computer to complete online interaction;
- No history of acute ischemic stroke, systemic lupus erythematosus, gastric ulcer, pneumonia, acute cardiac infarction, urinary tract infection, uterine fibroids, diabetes, osteoarthritis, or migraine.
- Able to understand and comply with study procedures and to provide written informed consent.
Exclusion Criteria:
- Currently or previously employed as a healthcare worker;
- Previously received systematic medical training;
- Currently involved in concurrent research that may interfere with the results of the present trial;
- The investigator considered that the participant had other conditions that might affect compliance or preclude participation.
Plan d'étude
Comment l'étude est-elle conçue ?
Détails de conception
- Objectif principal: Recherche sur les services de santé
- Répartition: Randomisé
- Modèle interventionnel: Affectation parallèle
- Masquage: Seul
Nombre de bras
Armes et Interventions
Groupe de participants / BrasGroupe de participants / Bras |
Intervention / TraitementIntervention / Traitement |
|---|---|
|
Expérimental: 3M-6D education GPT Group
Participants will first be trained in 3M-6D education, then use ChatGPT to complete a consultation task in unrestricted natural language in approximately 10 minutes.
|
Participants use ChatGPT to complete a standardized simulated clinical scenarios in unrestricted natural language.
3M-6D education is designed based on Cognitive Load Theory to reduce the cognitive burden on patients during medical interactions with AI and to improve the clarity and completeness of symptom reporting. Guided by cognitive load theory and the natural process physicians use to take medical histories, the investigators identified candidate information dimensions and developed a structured expression framework with six dimensions for public health queries through a Delphi expert consensus process. Participants were instructed to use the framework to describe their symptoms across these six dimensions; this process can typically be completed within three minutes, so the investigators call this approach three minutes six dimensions education (3M-6D education). |
|
Expérimental: 3M-6D education Gemini Group
Participants will first be trained in 3M-6D education, then use Gemini to complete a consultation task in unrestricted natural language in approximately 10 minutes.
|
Participants use Gemini to complete a standardized simulated clinical scenarios in unrestricted natural language.
3M-6D education is designed based on Cognitive Load Theory to reduce the cognitive burden on patients during medical interactions with AI and to improve the clarity and completeness of symptom reporting. Guided by cognitive load theory and the natural process physicians use to take medical histories, the investigators identified candidate information dimensions and developed a structured expression framework with six dimensions for public health queries through a Delphi expert consensus process. Participants were instructed to use the framework to describe their symptoms across these six dimensions; this process can typically be completed within three minutes, so the investigators call this approach three minutes six dimensions education (3M-6D education). |
|
Comparateur actif: GPT Group
Participants will use ChatGPT to complete a consultation task in unrestricted natural language in approximately 10 minutes.
|
Participants use ChatGPT to complete a standardized simulated clinical scenarios in unrestricted natural language.
|
|
Comparateur actif: Gemini Group
Participants will use Gemini to complete a consultation task in unrestricted natural language in approximately 10 minutes.
|
Participants use Gemini to complete a standardized simulated clinical scenarios in unrestricted natural language.
|
|
Aucune intervention: Control group
Participants will use non-AI tools such as internet searches and medical websites to complete a consultation task in unrestricted natural language in approximately 10 minutes.
|
Que mesure l'étude ?
Principaux critères de jugement
Principaux critères de jugement
Mesure des résultats |
Description de la mesure |
Délai |
|---|---|---|
|
Relevant conditions identification of the 3M-6D education GPT group compared with the GPT group
Délai: 1 hour.
|
Relevant conditions identification is defined as the proportion of participants whose final response includes the expert-defined final diagnosis or a relevant differential diagnosis.
|
1 hour.
|
|
Disposition concordance of the 3M-6D education GPT group compared with the GPT group
Délai: 1 hour.
|
Disposition concordance is defined as the proportion of participants whose final care recommendation matches the expert-defined level.
The five levels are self-care, routine outpatient care, urgent outpatient care, emergency department visit, and emergency medical services.
|
1 hour.
|
|
Relevant conditions identification of the 3M-6D education Gemini group compared with the Gemini group
Délai: 1 hour.
|
1 hour.
|
|
|
Disposition concordance of the 3M-6D education Gemini group compared with the Gemini group
Délai: 1 hour.
|
1 hour.
|
Mesures de résultats secondaires
Mesures de résultats secondaires
Mesure des résultats |
Description de la mesure |
Délai |
|---|---|---|
|
Relevant conditions identification of the 3M-6D education GPT group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Relevant conditions identification of the 3M-6D education Gemini group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Disposition concordance of the 3M-6D education GPT group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Disposition concordance of the 3M-6D education Gemini group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Red-flag identification in the 3M-6D education GPT group compared with the GPT group
Délai: 1 hour.
|
Red-flag identification is defined as the proportion of participants whose final response includes the key warning signs that experts defined for the assigned scenario.
|
1 hour.
|
|
Red-flag identification in the 3M-6D education GPT group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Red-flag identification in the 3M-6D education Gemini group compared with the Gemini group
Délai: 1 hour.
|
1 hour.
|
|
|
Red-flag identification in the 3M-6D education Gemini group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
NASA Task Load Index score of the 3M-6D education GPT group compared with the GPT group
Délai: 1 hour.
|
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician.
It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance.
Each domain is scored from 0 to 100.
The total score is the mean of the six domains.
Higher scores indicate greater perceived task load.
|
1 hour.
|
|
NASA Task Load Index score of the 3M-6D education GPT group compared with the control group
Délai: 1 hour.
|
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician.
It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance.
Each domain is scored from 0 to 100.
The total score is the mean of the six domains.
Higher scores indicate greater perceived task load.
|
1 hour.
|
|
NASA Task Load Index score of the 3M-6D education Gemini group compared with the Gemini group
Délai: 1 hour.
|
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician.
It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance.
Each domain is scored from 0 to 100.
The total score is the mean of the six domains.
Higher scores indicate greater perceived task load.
|
1 hour.
|
|
NASA Task Load Index score of the 3M-6D education Gemini group compared with the control group
Délai: 1 hour.
|
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician.
It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance.
Each domain is scored from 0 to 100.
The total score is the mean of the six domains.
Higher scores indicate greater perceived task load.
|
1 hour.
|
|
Relevant conditions identification of the 3M-6D education GPT group compared with the 3M-6D education Gemini group
Délai: 1 hour.
|
1 hour.
|
|
|
Disposition concordance of the 3M-6D education GPT group compared with the 3M-6D education Gemini group
Délai: 1 hour.
|
1 hour.
|
|
|
Red-flag identification in the 3M-6D education GPT group compared with the 3M-6D education Gemini group
Délai: 1 hour.
|
1 hour.
|
|
|
NASA Task Load Index score of the 3M-6D education GPT group compared with the 3M-6D education Gemini group
Délai: 1 hour.
|
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician.
It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance.
Each domain is scored from 0 to 100.
The total score is the mean of the six domains.
Higher scores indicate greater perceived task load.
|
1 hour.
|
Autres mesures de résultats
Autres mesures de résultats
Mesure des résultats |
Description de la mesure |
Délai |
|---|---|---|
|
Failure to identify red flags in the 3M-6D education GPT group compared with the GPT group
Délai: 1 hour.
|
Failure to identify red flags is defined as the proportion of participants whose final response does not include the expert-defined red-flag symptoms or warning signs for the assigned standardized simulated clinical scenario.
|
1 hour.
|
|
Failure to identify red flags in the 3M-6D education GPT group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Failure to identify red flags in the 3M-6D education Gemini group compared with the Gemini group
Délai: 1 hour.
|
1 hour.
|
|
|
Failure to identify red flags in the 3M-6D education Gemini group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Underestimation of disposition in the 3M-6D education GPT group compared with the GPT group
Délai: 1 hour.
|
Underestimation of disposition is defined as the proportion of participants whose final care recommendation is lower than the expert-defined disposition level for the assigned standardized simulated clinical scenario.
|
1 hour.
|
|
Underestimation of disposition in the 3M-6D education GPT group compared with the control group
Délai: 1 hour.
|
1 hour.
|
|
|
Underestimation of disposition in the 3M-6D education Gemini group compared with the Gemini group
Délai: 1 hour.
|
1 hour.
|
|
|
Underestimation of disposition in the 3M-6D education Gemini group compared with the control group
Délai: 1 hour.
|
1 hour.
|
Collaborateurs et enquêteurs
Parrainer
Parrainer
Collaborateurs
Collaborateurs
Dates d'enregistrement des études
Dates principales de l'étude
Début de l'étude (Réel)
Début de l'étude
Achèvement primaire (Réel)
Achèvement primaire
Achèvement de l'étude (Réel)
Achèvement de l'étude
Dates d'inscription aux études
Première soumission
Première soumission
Première soumission répondant aux critères de contrôle qualité
Première soumission répondant aux critères de contrôle qualité
Première publication (Réel)
Première publication
Mises à jour des dossiers d'étude
Dernière mise à jour publiée (Réel)
Dernière mise à jour publiée
Dernière mise à jour soumise répondant aux critères de contrôle qualité
Dernière mise à jour soumise répondant aux critères de contrôle qualité
Dernière vérification
Dernière vérification
Plus d'information
Termes liés à cette étude
Mots clés
Autres numéros d'identification d'étude
Autres numéros d'identification d'étude
- LAMP-1
Plan pour les données individuelles des participants (IPD)
Prévoyez-vous de partager les données individuelles des participants (DPI) ?
Informations sur les médicaments et les dispositifs, documents d'étude
Étudie un produit pharmaceutique réglementé par la FDA américaine
Étudie un produit d'appareil réglementé par la FDA américaine
Ces informations ont été extraites directement du site Web clinicaltrials.gov sans aucune modification. Si vous avez des demandes de modification, de suppression ou de mise à jour des détails de votre étude, veuillez contacter register@clinicaltrials.gov. Dès qu'un changement est mis en œuvre sur clinicaltrials.gov, il sera également mis à jour automatiquement sur notre site Web .