Smartphone AI Assistance for Prehospital ECG Interpretation (ECG-IA)

AI & Prehospital ECG Analysis: A Randomized Controlled Trial of a Smartphone Large Language Model for Occlusion Myocardial Infarction Detection by Prehospital Providers

Prehospital providers interpret 12-lead electrocardiograms (ECGs) under time pressure and without immediate expert support. Missed acute coronary occlusion - occlusion myocardial infarction (OMI) - delays reperfusion, while false positive interpretations trigger unnecessary catheterization laboratory activations. Multimodal large language models (LLMs) available on any smartphone can now analyze a photographed ECG, and prehospital providers have begun using them spontaneously. No randomized trial has evaluated whether this practice improves diagnostic performance.

This randomized controlled trial compares the diagnostic performance of prehospital providers interpreting ECG clinical vignettes with and without mandatory assistance from a single, version-locked smartphone large language model. Participants - paramedics, emergency medical technicians, nurses and physicians practicing in prehospital care in French-speaking Switzerland - are randomized 1:1 on a dedicated digital platform and answer 14 clinical vignettes presented in individually randomized order. Each vignette is built around a real, anonymized 12-lead ECG obtained during routine clinical care.

The primary outcome is the proportion of vignettes for which the participant correctly identifies the presence or absence of an OMI. Secondary outcomes are sensitivity, specificity, and the accuracy of the prehospital priority decision level.

Study Overview

Status

Recruiting

Conditions

Intervention / Treatment

Detailed Description

Design and setting. Two-arm parallel-group randomized controlled trial conducted on a dedicated digital platform. Phase 1 takes place during a single in-person session at the Swiss French-speaking prehospital clinical research conference (Morat, Switzerland, 2 September 2026). Phase 2 extends recruitment to prehospital emergency services across French-speaking Switzerland (September to November 2026) using an identical standardized protocol.

Randomization. Individual 1:1 allocation performed automatically by the platform using permuted blocks of variable size (4 to 6), without stratification, after the demographic questionnaire and before the first vignette. Allocation cannot be changed once assigned. The presentation order of the 14 vignettes is independently randomized for each participant, which neutralizes position effects and prevents copying between neighbouring participants.

Intervention. Participants allocated to the intervention arm must consult the study-imposed large language model for every vignette before submitting their answer; the platform locks the submit button until use of the tool is confirmed. A single model (OpenAI GPT-4o) in an API version locked for the entire study is used by all participants, with a standardized, non-modifiable prompt identical for every participant and every vignette. Participants cannot add text, ask follow-up questions, or provide additional clinical context. Exposure to the tool is mandatory, but adherence to its interpretation is not: participants remain free to base their final answer on their own clinical reasoning. Control participants interpret the same ECGs unaided, with smartphones turned face down and out of reach.

Reference standard. For each vignette, the expected answers were defined a priori by the study cardiologist and locked, with any subsequent modification time-stamped in the platform audit log. The reference standard is the answer expected of a prehospital provider at the point of care, anchored on coronary angiography wherever angiography is discriminant. For non-ischaemic mimics, the expected answer depends on whether the acute presentation allows the condition to be distinguished from a coronary occlusion: it does not for the Takotsubo case included in the study, whose expected answer is therefore "yes", whereas acute pericarditis, being usually recognizable, has an expected answer of "no". This pre-specified departure from a purely angiographic standard is reported as such, and the primary analysis is repeated in a sensitivity analysis in which all mimics are classified as non-OMI.

Blinding. Participants and investigators cannot be masked. The statistician conducting the primary analysis receives coded group labels only, and the allocation key is released after database lock and approval of the statistical analysis plan.

Statistical analysis. Mixed-effects logistic regression with correct OMI identification at vignette level as the dependent variable, arm as the main fixed effect, grouped ECG category as a fixed covariate, and random intercepts for participant and for vignette. Intention-to-treat, alpha 0.05 two-sided. The target is 130 evaluable participants (65 per arm), giving 80% power to detect an absolute increase from 0.65 to 0.80 with an intraclass correlation up to approximately 0.43. Allowing for approximately 10% of sessions to be abandoned before completion, planned enrollment is 144 (72 per arm).

ECG material. All tracings originate from routine clinical care at the Geneva University Hospitals, were selected by a cardiologist to cover predefined electrocardiographic categories, and were anonymized at source before transmission to the research team. No patient is enrolled, followed or contacted, and the research team holds no key allowing re-identification.

Pilot. A technical pilot involving a small number of prehospital providers was conducted before the study start date to test the platform. Pilot sessions took place outside the standardized study conditions and are identified in the database by a dedicated centre code; their data are excluded from all analyses, as pre-specified in the protocol.

Study Type

Interventional

Enrollment (Estimated)

144

Phase

  • Not Applicable

Contacts and Locations

This section provides the contact details for those conducting the study, and information on where this study is being conducted.

Study Contact

Study Locations

    • Canton of Fribourg
      • Murten/Morat, Canton of Fribourg, Switzerland, 3280
        • Recruiting
        • Caserne des pompiers de Morat (Swiss French-speaking prehospital clinical research conference)
        • Contact:

Participation Criteria

Researchers look for people who fit a certain description, called eligibility criteria. Some examples of these criteria are a person's general health condition or prior treatments.

Eligibility Criteria

Ages Eligible for Study

  • Adult
  • Older Adult

Accepts Healthy Volunteers

Yes

Description

Inclusion Criteria:

  • Prehospital care provider practising in French-speaking Switzerland
  • Any level of training: emergency medical technician, paramedic (ES), nurse (ES/HES) in prehospital emergency care, or prehospital emergency physician
  • Electronic informed consent signed before randomization

Exclusion Criteria:

  • Cardiologist
  • Any person not practising in prehospital care
  • Insufficient command of written French to answer the vignettes reliably
  • Refusal to participate or withdrawal of consent

Study Plan

This section provides details of the study plan, including how the study is designed and what the study is measuring.

How is the study designed?

Design Details

  • Primary Purpose: Diagnostic
  • Allocation: Randomized
  • Interventional Model: Parallel Assignment
  • Masking: Single

Arms and Interventions

Participant Group / Arm
Intervention / Treatment
No Intervention: Control: unaided ECG interpretation
Participants interpret each of the 14 ECG vignettes without any assistance. Smartphones are turned face down and out of reach for the duration of the session. No intervention is administered.
Experimental: AI-assisted ECG interpretation
Participants must consult the study-imposed large language model for every vignette before submitting their answer. The platform locks the submit button until use of the tool is confirmed. Participants remain free not to follow the interpretation produced by the model and may base their final answer on their own clinical reasoning.
The platform transmits the ECG image to a single large language model (OpenAI GPT-4o, API snapshot gpt-4o-2024-08-06), locked for the entire study, together with a standardised prompt identical for all participants and all vignettes: "I am on an urgent prehospital call with a patient who presents this ECG. Analyse it and tell me what you think." Participants cannot modify the prompt, ask follow-up questions or provide additional clinical context. The model version and system fingerprint returned by the API are recorded for every call. The model's interpretation is displayed within the vignette. Use of the tool is mandatory; adherence to its interpretation is not.
Other Names:
  • ChatGPT
  • gpt-4o-2024-08-06

What is the study measuring?

Primary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Proportion of vignettes with correct identification of occlusion myocardial infarction (OMI) status
Time Frame: Single study session, approximately 90 minutes; 14 vignettes per participant
For each of the 14 vignettes, the participant answers a binary question: "At this stage of care, is an OMI (acute coronary occlusion) likely? Yes / No". Responses are scored against a reference standard defined a priori, vignette by vignette, by the study cardiologist and locked before data collection. This reference standard is the answer expected of a prehospital provider at the point of care, anchored on coronary angiography wherever angiography is discriminant. For non-ischaemic mimics the expected answer depends on whether the acute presentation allows the condition to be distinguished from a coronary occlusion: it does not for the Takotsubo case (expected answer "yes"), whereas acute pericarditis is usually recognisable (expected answer "no"). This pre-specified departure from a purely angiographic standard is reported as such, and the analysis is repeated in a sensitivity analysis classifying all mimics as non-OMI. The outcome is the proportion of correctly classified vignettes
Single study session, approximately 90 minutes; 14 vignettes per participant

Secondary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Sensitivity of OMI detection
Time Frame: Single study session, approximately 90 minutes
Proportion of correctly identified OMI among vignettes whose expected answer is "yes", compared between arms. Analysed by mixed-effects logistic regression restricted to these vignettes, with arm as fixed effect and a random intercept per participant.
Single study session, approximately 90 minutes
Specificity of OMI detection
Time Frame: Single study session, approximately 90 minutes
Proportion of correctly identified non-OMI among vignettes whose expected answer is "no" (non-occlusive ischaemia, indeterminate ischaemia, recognisable non-ischaemic mimics, normal or non-specific ECGs), compared between arms. Same modelling approach as for sensitivity.
Single study session, approximately 90 minutes
Accuracy of the prehospital priority decision level
Time Frame: Single study session, approximately 90 minutes
Proportion of vignettes for which the participant selects the correct prehospital priority decision level, on a pre-specified five-level scale ranging from high-risk acute coronary syndrome with activation of the STEMI pathway to clearly non-cardiac symptoms. The expected level is defined a priori for each vignette and locked before data collection.
Single study session, approximately 90 minutes

Other Outcome Measures

Outcome Measure
Measure Description
Time Frame
Diagnostic accuracy on the closed-list diagnosis question
Time Frame: Single study session, approximately 90 minutes
For each vignette the participant selects a single diagnosis from a closed list of 86 items grouped in 8 families, using a filterable selector. Two variables are produced automatically: exact match with the expected diagnosis, and match with the expected diagnosis or with any diagnosis pre-specified as clinically acceptable. Reported descriptively by ECG category and by arm, then modelled by mixed-effects logistic regression if numbers allow. The primary analysis uses the exact-or-acceptable variable; the strict variable is reported as a sensitivity analysis.
Single study session, approximately 90 minutes
Self-reported confidence and calibration
Time Frame: Single study session, approximately 90 minutes
Confidence in the answer is reported by the participant for each vignette on a five-point Likert scale of self-reported diagnostic confidence, ranging from 1 (not at all confident) to 5 (very confident). A higher score indicates greater self-reported confidence; confidence is a descriptive measure and a higher score is not in itself a better outcome. Confidence is compared between arms by mixed-effects ordinal logistic regression. Calibration is assessed as the concordance between reported confidence and the actual correctness of the answer, using calibration curves and the Brier score, which ranges from 0 to 1, a lower score indicating better calibration.
Single study session, approximately 90 minutes
Self-reported influence of the AI on the final answer
Time Frame: Single study session, approximately 90 minutes
Degree to which the participant reports having been influenced by the model's interpretation, on a five-point Likert scale, collected for each vignette in the intervention arm only. Analysed descriptively and in relation to the correctness of the answer. No comparison between AI tools is conducted, the model version being fixed.
Single study session, approximately 90 minutes
Response time per vignette
Time Frame: Single study session, approximately 90 minutes
Time from display of the vignette to submission of the answer, measured automatically in seconds by the platform. Analysed by linear mixed-effects model on the log of the response time, with arm as fixed effect and a random intercept per participant; a sensitivity analysis compares individual median times between arms. A longer time is expected in the intervention arm owing to consultation of the model, and is treated as a feasibility outcome for prehospital use.
Single study session, approximately 90 minutes
Diagnostic performance by level of training
Time Frame: Single study session, approximately 90 minutes
Diagnostic performance according to the participant's level of training (emergency medical technician, paramedic, nurse, physician), and interaction between arm and level of training. The study is not powered for this comparison and the distribution of training levels depends on spontaneous recruitment; results are hypothesis-generating and reported with confidence intervals only. A minimum of 10 participants per category is required for separate analysis, failing which categories are pooled according to a rule documented before unblinding.
Single study session, approximately 90 minutes

Collaborators and Investigators

This is where you will find people and organizations involved with this study.

Sponsor

Collaborators

Study record dates

These dates track the progress of study record and summary results submissions to ClinicalTrials.gov. Study records and reported results are reviewed by the National Library of Medicine (NLM) to make sure they meet specific quality control standards before being posted on the public website.

Study Major Dates

Study Start (Actual)

September 2, 2026

Primary Completion (Estimated)

November 30, 2026

Study Completion (Estimated)

November 30, 2026

Study Registration Dates

First Submitted

August 27, 2026

First Submitted That Met QC Criteria

September 4, 2026

First Posted (Actual)

September 9, 2026

Study Record Updates

Last Update Posted (Actual)

September 9, 2026

Last Update Submitted That Met QC Criteria

September 4, 2026

Last Verified

September 1, 2026

More Information

Terms related to this study

Other Study ID Numbers

  • ECG-IA-2026
  • 2026-01430 (Other Identifier: Cantonal Research Ethics Committee of Geneva (CCER) - BASEC)

Plan for Individual participant data (IPD)

Plan to Share Individual Participant Data (IPD)?

YES

IPD Plan Description

Fully anonymised individual participant data - responses to all vignettes, demographic variables and platform-generated technical variables - will be deposited in an open repository (Zenodo or OSF) at the time of publication, in accordance with FAIR principles. The ECG tracings themselves are excluded from this deposit: they originate from routine clinical care at the Geneva University Hospitals, which authorise their use for this study only and do not permit transmission to third parties. Requests concerning the tracings must be addressed to the Geneva University Hospitals.

IPD Sharing Time Frame

From the date of publication, with no end date.

IPD Sharing Access Criteria

Open access, no restriction, no request procedure.

IPD Sharing Supporting Information Type

  • STUDY_PROTOCOL
  • SAP
  • ICF
  • ANALYTIC_CODE

Drug and device information, study documents

Studies a U.S. FDA-regulated drug product

No

Studies a U.S. FDA-regulated device product

No

This information was retrieved directly from the website clinicaltrials.gov without any changes. If you have any requests to change, remove or update your study details, please contact register@clinicaltrials.gov. As soon as a change is implemented on clinicaltrials.gov, this will be updated automatically on our website as well.