AI in the ED Cologne (KINA-CO)

July 21, 2026 updated by: Volker Burst, University of Cologne

Retrospective Validation of Large Language Models (LLM) for the Prognostic Assessment of Clinical Parameters in Emergency Department and Evaluation of the Impact of Automated Anonymisation Methods

This retrospective, non-interventional study evaluates the prognostic performance of open-weight Large Language Models (LLMs) in the setting of a German academic emergency department. Using a full census of all consecutive emergency department cases at University Hospital Cologne between 01 January 2023 and 31 December 2025 (approximately 100,000 cases), the study assesses whether LLMs can make reliable prognostic predictions (e.g., hospital admission, imaging, diagnosis, placement) based on the initial history, vital signs, and triage category. In addition, it quantifies how strongly automated anonymization and perturbation procedures affect the models' diagnostic accuracy. This is an Investigator-Initiated Trial (IIT) with no intervention on patients.

Study Overview

Status

Active, not recruiting

Detailed Description

The study analyzes a retrospective cohort of all emergency department cases at the Central Emergency Department of University Hospital Cologne (01 January 2023 - 31 December 2025). Data originate from the hospital information system (HIS) and are provided in pseudonymized form via the Medical Data Integration Center (MeDIC) of University Hospital Cologne, acting as an independent trusted third party. Extracted data include sociodemographic data (age/year of birth, sex), clinical vital signs (blood pressure, heart rate, respiratory rate, oxygen saturation, temperature, GCS), medical free text (triage records, physician history and admission findings), and process/outcome data serving as the gold standard (ICD-10 diagnoses, imaging performed, admission status, timestamps/length of stay). LLM processing takes place on premise on local compute clusters or in a contractually secured enterprise environment with zero data retention; open-weight models are used. Two arms are compared: Arm A (original data) vs. Arm B (anonymized/perturbed/synthesized data). The primary endpoint is diagnostic accuracy (AUROC, F1 score) against the documented clinical outcome. Working hypotheses: (1) modern LLMs are non-inferior to the human assessment (non-inferiority); (2) modern anonymization procedures reduce model performance by less than 5% (relative performance loss).

Study Type

Observational

Enrollment (Estimated)

100000

Contacts and Locations

This section provides the contact details for those conducting the study, and information on where this study is being conducted.

Study Locations

      • Cologne, Germany, 50937
        • Department of Internal Medicine II, University Hospital Cologne

Participation Criteria

Researchers look for people who fit a certain description, called eligibility criteria. Some examples of these criteria are a person's general health condition or prior treatments.

Eligibility Criteria

Ages Eligible for Study

  • Child
  • Adult
  • Older Adult

Accepts Healthy Volunteers

No

Sampling Method

Non-Probability Sample

Study Population

All consecutive treatment cases at the Central Emergency Department of University Hospital Cologne during 01 January 2023 - 31 December 2025 (full census, approximately 100,000 cases).

Description

Inclusion Criteria:

  • All consecutive treatment cases at the Central Emergency Department of University Hospital Cologne during the period 01 January 2023 - 31 December 2025 (full census, consecutive inclusion).

Exclusion Criteria:

  • Documented objection to the scientific use of the data pursuant to Art. 21 General data protection Regulation (GDPR).
  • Cases lacking the minimum data required for analysis (triage/history and documented outcome).

Study Plan

This section provides details of the study plan, including how the study is designed and what the study is measuring.

How is the study designed?

Design Details

Cohorts and Interventions

Group / Cohort
Analytic Arm A
Original data will be used for the analysis
Analytic Arm B
Anonymized/perturbed data will be used for the analysis

What is the study measuring?

Primary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Diagnostic accuracy of the LLM predictions (AUROC, F1 score) compared with the clinical gold standard
Time Frame: From enrollment to the end of retrospective observation period at 1 year
assessment at the level of the individual emergency department encounter).]
From enrollment to the end of retrospective observation period at 1 year

Secondary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Relative performance loss of the models between original data (Arm A) and anonymized/perturbed data (Arm B); hypothesis < 5%.
Time Frame: From enrollment to the end of retrospective observation period at 1 year
Relative loss of diagnostic accuracy (AUROC, F1 score) when the models are applied to anonymized/perturbed data (Arm B) compared with original data (Arm A), expressed as the relative percentage change. Non-inferiority is assumed if the relative performance loss is below 5%.]
From enrollment to the end of retrospective observation period at 1 year
Sensitivity, specificity, PPV/NPV for binary endpoints and agreement of the triage assessment (Cohen's kappa / Krippendorff's alpha).
Time Frame: From enrollment to the end of retrospective observation period at 1 year
Agreement between LLM output and the documented reference for binary endpoints (e.g., admission yes/no), reported as sensitivity, specificity, and positive/negative predictive value; agreement on the ordinal triage category is reported using Cohen's kappa or Krippendorff's alpha
From enrollment to the end of retrospective observation period at 1 year

Collaborators and Investigators

This is where you will find people and organizations involved with this study.

Investigators

  • Principal Investigator: Volker Burst, Prof., Department of Internal Medicine II, University Hospital Cologne

Study record dates

These dates track the progress of study record and summary results submissions to ClinicalTrials.gov. Study records and reported results are reviewed by the National Library of Medicine (NLM) to make sure they meet specific quality control standards before being posted on the public website.

Study Major Dates

Study Start (Actual)

July 1, 2026

Primary Completion (Estimated)

July 31, 2027

Study Completion (Estimated)

December 31, 2027

Study Registration Dates

First Submitted

July 21, 2026

First Submitted That Met QC Criteria

July 21, 2026

First Posted (Actual)

July 24, 2026

Study Record Updates

Last Update Posted (Actual)

July 24, 2026

Last Update Submitted That Met QC Criteria

July 21, 2026

Last Verified

July 1, 2026

More Information

Terms related to this study

Other Study ID Numbers

  • KINA-CO
  • 26-1064 (Other Identifier: Ethics Committee)

Plan for Individual participant data (IPD)

Plan to Share Individual Participant Data (IPD)?

NO

Drug and device information, study documents

Studies a U.S. FDA-regulated drug product

No

Studies a U.S. FDA-regulated device product

No

This information was retrieved directly from the website clinicaltrials.gov without any changes. If you have any requests to change, remove or update your study details, please contact register@clinicaltrials.gov. As soon as a change is implemented on clinicaltrials.gov, this will be updated automatically on our website as well.

Clinical Trials on Triage

Subscribe