Tämä sivu käännettiin automaattisesti, eikä käännösten tarkkuutta voida taata. Katso englanninkielinen versio lähdetekstiä varten.

Improving AI-Assisted Medical Diagnosis and Triage by the General Public

keskiviikko 16. syyskuuta 2026 päivittänyt: Ihsan Ayyub Qazi, PhD, Lahore University of Management Sciences
This study is a randomized controlled trial (RCT) investigating whether access to a new LLM interface can improve medical triage and diagnostic accuracy for laypeople compared to access to a standard LLM interface. It addresses previous findings where laypeople using standard LLMs performed worse than those using conventional methods (e.g., web search) due to incomplete symptom sharing and poor interpretation of AI advice. To address this, the research tests a structured LLM system that proactively asks clinical history questions before providing a standardized, easy-to-read diagnostic output.

Tutkimuksen yleiskatsaus

Tila

Rekrytointi

Ehdot

Interventio / Hoito

Yksityiskohtainen kuvaus

As large language models (LLMs) become widely accessible, a growing share of the public already turns to AI-powered chatbots for health-related information: surveys suggest that one in six American adults consults AI chatbots for health queries at least once a month. At the same time, diagnostic errors and misplaced triage decisions represent a persistent source of preventable patient harm globally [4]. This has prompted considerable interest in whether LLMs can serve as a reliable "front door" to the healthcare system for patients who lack immediate access to a clinician.

However, incidents involving misguided medical suggestions from LLMs to the general public such as fatal overdose and erroneous diagnosis leading to life-threatening treatment delay have been reported. A recent controlled study shows LLMs perform significantly worse for medical assistance in real-world settings compared to their performance on controlled benchmarks. Bean et al. [1] conducted a large preregistered study with 1,298 UK participants in which laypeople were assigned to receive assistance from one of three popular LLMs (GPT-4o, Llama 3, or Command R+) or to use a source of their choice when assessing ten standardized medical scenarios. Although the LLMs alone correctly identified relevant conditions in up to 94.9% of cases, participants using those same LLMs did so in fewer than 34.5% of cases, significantly worse than the group using their typical home resources (∼60%). Similarly, triage accuracy showed no signifcant difference between LLM users and controls, with an overall correct response rate of 43.0% across all groups. This poor performance can be linked to two main failure modes: users providing incomplete symptom information to the AI, and users failing to correctly interpret or act upon the LLM's output.

To address these failure modes, the study tests a new GPT-4o interface wrapped with a fixed system prompt. Rather than passively responding to whatever the user volunteers, the AI is instructed to ask clarifying questions to gather a proper clinical history before offering any medical suggestions. Once the model judges that enough information is gathered, it provides a structured, easy-to-read response detailing possible conditions, their likelihoods, and a clear triage recommendation.

The trial is structured as a two-arm, single-blind RCT. Both arms retain access to whatever assistance methods participants would normally use at home (e.g., web search), and both arms additionally get a GPT-4o-based LLM interface. The difference lies in which interface: the treatment arm receives the new interface described above, while the control arm receives a standard, unmodified GPT-4o interface with no special prompting. Participants are restricted to using the provided LLM interface (GPT-4o) only, and not other LLMs. AI-overview in web searches will be disabled through an extension. The trial utilizes ten previously validated clinical vignettes covering various medical urgencies, ranging from self-care routines up to ambulance-level emergencies, with each participant randomly assigned two of the ten. To achieve sufficient statistical power (accounting for within-participant clustering across those two responses), the target sample size is calculated at 220 total participants, split evenly with 110 individuals per arm.

The participant pool is drawn from the enrolled students and administrative staff at the Lahore University of Management Sciences (LUMS) in Pakistan. To ensure the sample consists strictly of laypeople, anyone currently enrolled in or who has completed a medical, nursing, or allied health professional degree is explicitly excluded from participating.

The study is driven by two co-primary hypotheses: participants with access to the new LLM interface will identify relevant medical conditions at a higher rate than participants using the standard LLM interface, and they will correctly assess the urgency of the medical scenarios at a higher rate. Accuracy is measured by comparing participant responses to a physician-generated gold-standard list using fuzzy matching, then analyzed via mixed-effects logistic regression with participant-level covariates and a Bonferroni-corrected significance threshold. The study also tracks secondary outcomes - self-reported confidence and time spent per scenario - and several exploratory analyses, including an LLM-alone benchmark, moderation by LLM experience, internet usage patterns, digital literacy, and scenario-level variation in accuracy.

Opintotyyppi

Interventio

Ilmoittautuminen (Arvioitu)

220

Vaihe

  • Ei sovellettavissa

Yhteystiedot ja paikat

Tässä osiossa on tutkimuksen suorittajien yhteystiedot ja tiedot siitä, missä tämä tutkimus suoritetaan.

Opiskeluyhteys

Tutki yhteystietojen varmuuskopiointi

Opiskelupaikat

    • Punjab Province
      • Lahore, Punjab Province, Pakistan, 54000
        • Rekrytointi
        • Lahore University of Management Sciences
        • Ottaa yhteyttä:
        • Ottaa yhteyttä:

Osallistumiskriteerit

Tutkijat etsivät ihmisiä, jotka sopivat tiettyyn kuvaukseen, jota kutsutaan kelpoisuuskriteereiksi. Joitakin esimerkkejä näistä kriteereistä ovat henkilön yleinen terveydentila tai aiemmat hoidot.

Kelpoisuusvaatimukset

Opintokelpoiset iät

  • Aikuinen
  • Vanhempi Aikuinen

Hyväksyy terveitä vapaaehtoisia

Joo

Kuvaus

Inclusion Criteria:

  • Enrolled student or employed administrative staff at LUMS.
  • 18 years of age or older.
  • Able to read and understand English.

Exclusion Criteria:

  • Individuals with any formal education or professional training in medicine, nursing, or any other healthcare profession.

Opintosuunnitelma

Tässä osiossa on tietoja tutkimussuunnitelmasta, mukaan lukien kuinka tutkimus on suunniteltu ja mitä tutkimuksella mitataan.

Miten tutkimus on suunniteltu?

Suunnittelun yksityiskohdat

  • Ensisijainen käyttötarkoitus: Diagnostiikka
  • Jako: Satunnaistettu
  • Inventiomalli: Rinnakkaistehtävä
  • Naamiointi: Yksittäinen

Aseet ja interventiot

Osallistujaryhmä / Arm
Interventio / Hoito
Kokeellinen: Treatment Arm (new GPT-4o Interface)
Participants can access any assistance methods they would typically employ (e.g., web search or health portals) in addition to a new LLM interface (based on GPT-4o) to complete medical scenarios. The new LLM interface uses a fixed system prompt that (a) instructs the model to ask targeted clarifying questions before providing any diagnostic or triage suggestions, and (b) requires all final responses to follow a structured template listing: possible conditions, approximate likelihood of each, and a recommended triage with brief reasoning.
Participants can access any assistance methods they would typically employ (e.g., web search or health portals) in addition to a new LLM (GPT-4o) interface to complete medical scenarios. The new LLM interface uses a fixed system prompt that (a) instructs the model to ask targeted clarifying questions before providing any diagnostic or triage suggestions, and (b) requires all final responses to follow a structured template listing: possible conditions, approximate likelihood of each, and a recommended triage with brief reasoning.
Placebo Comparator: Control
Participants can use any assistance methods they would typically employ (e.g., web search or health portals) in addition to a standard LLM (GPT-4o) to complete medical scenarios. AI-overview in web searches will be disabled via an extension. They would not be allowed to access any LLMs other than the standard LLM interface.
Participants use any assistance methods they would typically employ at home (e.g., Google or health portals) in addition to a standard LLM (GPT-4o) to complete medical scenarios. AI-overview in web searches will be disabled via an extension. They would not be allowed to access any LLMs other than the standard LLM interface.

Mitä tutkimuksessa mitataan?

Ensisijaiset tulostoimenpiteet

Tulosmittaus
Toimenpiteen kuvaus
Aikaikkuna
Urgency Assessment Accuracy
Aikaikkuna: Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.
The primary outcome will be the percentage of correct urgency assessments, ranging from 0 to 100%, where higher scores indicate better urgency assessment (or triage) performance. The rating which participants give for the urgency of each case, will be measured on a five-point scale: Self-care, Routine GP, Urgent Primary Care, Accident & Emergency, and Ambulance. Responses will be compared against the gold-standard answers to produce an accuracy measure. The primary outcome will be compared at the case-level between the randomized groups.
Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.
Condition Identification Accuracy
Aikaikkuna: Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.
The co-primary outcome is the Condition Identification Accuracy, which is the percentage of cases in which the condition was correctly identified, ranging from 0 to 100%., where higher scores indicate better medical condition identification performance. Participants name all medical conditions they considered relevant to their decision. A response is scored as correct for that scenario if at least one named condition matches the physician-generated gold-standard list of relevant conditions. The co-primary outcome will be compared at the case-level between the randomized groups.
Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.

Toissijaiset tulostoimenpiteet

Tulosmittaus
Toimenpiteen kuvaus
Aikaikkuna
Self-Reported Confidence
Aikaikkuna: Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.
Participants' self-reported confidence in their urgency assessments, measured using a scale ranging from 0 to 100.
Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.
Time Spent
Aikaikkuna: Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.
The amount of time participants spend completing each medical scenario in milliseconds.
Assessed at a single time point for each case, during the scheduled diagnostic evaluation session, which takes place between 0-5 days after participant enrollment.

Yhteistyökumppanit ja tutkijat

Täältä löydät tähän tutkimukseen osallistuvat ihmiset ja organisaatiot.

Sponsori

Tutkijat

  • Päätutkija: Ayesha Ali, PhD, Lahore University of Management Sciences (LUMS)
  • Opintojen puheenjohtaja: Ihsan Ayyub Qazi, PhD, Lahore University of Management Sciences (LUMS)
  • Päätutkija: Zafar Ayyub Qazi, PhD, Lahore University of Management Sciences (LUMS)

Opintojen ennätyspäivät

Nämä päivämäärät seuraavat ClinicalTrials.gov-sivustolle lähetettyjen tutkimustietueiden ja yhteenvetojen edistymistä. National Library of Medicine (NLM) tarkistaa tutkimustiedot ja raportoidut tulokset varmistaakseen, että ne täyttävät tietyt laadunvalvontastandardit, ennen kuin ne julkaistaan ​​julkisella verkkosivustolla.

Opi tärkeimmät päivämäärät

Opiskelun aloitus (Todellinen)

Keskiviikko 1. heinäkuuta 2026

Ensisijainen valmistuminen (Arvioitu)

Torstai 1. heinäkuuta 2027

Opintojen valmistuminen (Arvioitu)

Torstai 1. heinäkuuta 2027

Opintoihin ilmoittautumispäivät

Ensimmäinen lähetetty

Keskiviikko 22. heinäkuuta 2026

Ensimmäinen toimitettu, joka täytti QC-kriteerit

Keskiviikko 22. heinäkuuta 2026

Ensimmäinen Lähetetty (Todellinen)

Maanantai 27. heinäkuuta 2026

Tutkimustietojen päivitykset

Viimeisin päivitys julkaistu (Todellinen)

Torstai 17. syyskuuta 2026

Viimeisin lähetetty päivitys, joka täytti QC-kriteerit

Keskiviikko 16. syyskuuta 2026

Viimeksi vahvistettu

Tiistai 1. syyskuuta 2026

Lisää tietoa

Tähän tutkimukseen liittyvät termit

Avainsanat

Muut tutkimustunnusnumerot

  • AI-Assisted-Diag-Public789

Yksittäisten osallistujien tietojen suunnitelma (IPD)

Aiotko jakaa yksittäisten osallistujien tietoja (IPD)?

EI

Lääke- ja laitetiedot, tutkimusasiakirjat

Tutkii yhdysvaltalaista FDA sääntelemää lääkevalmistetta

Ei

Tutkii yhdysvaltalaista FDA sääntelemää laitetuotetta

Ei

Nämä tiedot haettiin suoraan verkkosivustolta clinicaltrials.gov ilman muutoksia. Jos sinulla on pyyntöjä muuttaa, poistaa tai päivittää tutkimustietojasi, ota yhteyttä register@clinicaltrials.gov. Heti kun muutos on otettu käyttöön osoitteessa clinicaltrials.gov, se päivitetään automaattisesti myös verkkosivustollemme .