- ICH GCP
- US Clinical Trials Registry
- Clinical Trial NCT06814327
Assessment of Organ Failure Risk Predictions in ICU (AI4ICU-Obs)
August 27, 2025 updated by: ETH Zurich
Prospective Assessment of Risk Predictions of Organ Failure in the Intensive Care Unit - Comparing Accuracy of Human and AI Risk Predictions
During this observational study, the investigators aim to assess the ability of ICU clinicians to predict the risk of impending organ failure and retrospectively compare it to the performance of previously published machine learning models.
The central hypothesis of this study is that the treating physician can predict impending organ failure in adult ICU patients with similar accuracy as the best previously publishes machine learning models.
Study Overview
Status
Active, not recruiting
Conditions
Detailed Description
In this observational study, clinician's (physicians and nurses) assessment of the estimated imminent organ failure risk in an ICU setting are prospectively collected.
Circulatory failure is investigated in the primary objective, and respiratory failure, renal failure, and mortality are investigated in secondary objectives.
These assessments investigate the predictive performance and influencing factors for clinician prediction.
The assessments will be collected in questionnaires and be performed by the clinicians directly involved in the patient treatment and by clinicians who are not actively responsible for the patient treatment.
Furthermore, this study aims to benchmark these risk assessments made by healthcare professionals against retrospectively generated AI risk scores for the same patients and timepoints.
The AI risk scores will be calculated retrospectively from a set of models from a systematic search of the current literature.
The AI models that will be employed for this analysis will be identified as indicated by a systematic review protocol and must satisfy the following two criteria: they do not require any data beyond what is routinely collected during an ICU stay and may be accessed as open source.
Such a comparison is vital for the understanding of the relative accuracy and reliability of AI-based predictions in the context of organ failure risk compared to human performance.
The data and findings from this study are anticipated to provide evidence for the clinical utility of AI-based risk scores and pave the way for future research into the optimization of AI systems for healthcare applications.
Study Type
Observational
Enrollment (Actual)
499
Contacts and Locations
This section provides the contact details for those conducting the study, and information on where this study is being conducted.
Study Locations
-
-
Canton of Bern
-
Bern, Canton of Bern, Switzerland, 3010
- University Hospital Inselspital, Berne
-
-
Participation Criteria
Researchers look for people who fit a certain description, called eligibility criteria. Some examples of these criteria are a person's general health condition or prior treatments.
Eligibility Criteria
Ages Eligible for Study
- Adult
- Older Adult
Accepts Healthy Volunteers
No
Sampling Method
Probability Sample
Study Population
Adult Intensive Care Unit at the University Hospital Bern
Description
Inclusion Criteria:
- patient minimum age of 18 years
- emergency admission to the ICU
- arterial line in place
Exclusion Criteria:
- documented refusal (on the general consent form) to participate to clinical research
- patients with neurologic conditions that impair the patient's level of consciousness (including, but not limited to stroke, traumatic brain injury, intracranial hemorrhage, CNS infections; except polytrauma)
- patients on mechanical circulatory support systems (IABP, VA-ECMO, Impella, VAD) or extracorporeal membrane oxygenation (VV-ECMO) at any time during their ICU stay;
- patients receiving end-of-life care or are admitted for the sole purpose of evaluating organ donation
Study Plan
This section provides details of the study plan, including how the study is designed and what the study is measuring.
How is the study designed?
Design Details
Cohorts and Interventions
Group / Cohort |
|---|
|
Adult ICU patients
|
What is the study measuring?
Primary Outcome Measures
Outcome Measure |
Measure Description |
Time Frame |
|---|---|---|
|
Clinician prediction of circulatory failure within 8 hours compared to published ML models
Time Frame: Assessments are collected within the first 72 hours following admission.
|
This outcome compares the area under the receiver operating characteristic curve (auROC) for two methods of predicting circulatory failure within 8 hours of each assessment time point: (1) ICU clinicians' risk estimates, and (2) previously published machine learning (ML) models applied retrospectively.
For each assessment, we compute the auROC separately for clinicians and for the ML model for the same time points and patients.
The difference in auROC (clinician minus ML) is the main measure of interest, evaluated under a non-inferiority framework with a margin of 0.025.
|
Assessments are collected within the first 72 hours following admission.
|
Secondary Outcome Measures
Outcome Measure |
Measure Description |
Time Frame |
|---|---|---|
|
Clinician prediction of respiratory failure within 24 hours compared to published ML models
Time Frame: Assessments are collected within the first 72 hours following admission.
|
This outcome compares the area under the receiver operating characteristic curve (auROC) for two methods of predicting respiratory failure within 24 hours of each assessment time point: (1) ICU clinicians' risk estimates, and (2) previously published machine learning (ML) models applied retrospectively.
For each assessment, we compute the auROC separately for clinicians and for the ML model.
The difference in auROCs (clinician minus ML) is the main measure of interest, evaluated using the same methodological framework as the primary outcome.
|
Assessments are collected within the first 72 hours following admission.
|
Other Outcome Measures
Outcome Measure |
Measure Description |
Time Frame |
|---|---|---|
|
Exploratory analysis of clinician prediction of renal failure within 48 hours compared to published ML models
Time Frame: Assessments are collected within the first 72 hours following admission.
|
This exploratory outcome compares the area under the receiver operating characteristic curve (auROC) for two methods of predicting renal failure within 48 hours of each assessment time point: (1) ICU clinicians' risk estimates, and (2) previously published machine learning (ML) models applied retrospectively.
For each assessment, we compute the auROC separately for clinicians and for the ML model.
The difference in auROCs is evaluated using the same methodological framework as the primary outcome.
|
Assessments are collected within the first 72 hours following admission.
|
|
Exploratory analysis of clinician mortality prediction compared to published ML models
Time Frame: Assessments are collected within the first 72 hours following admission.
|
All-cause mortality prediction by McNemar's test.
Evaluated and tested using the performance of respective assessments by clinicians (binary response) and corresponding (paired) predictions by a machine learning model (probability prediction reduced to a binary response) trained on historical data for the three individual binary outcomes of all-cause mortality <28-day, <6-months, <12-months.
For each horizon separately the machine-based probabilities are thresholded to match the sensitivity of the clinicians, a McNemar's test is performed to test for a significant difference in predictive capabilities and p-values will be adjusted to account for multiple testing if necessary.
|
Assessments are collected within the first 72 hours following admission.
|
|
Exploratory analysis of predictive accuracy of treating physicians versus treating nurses
Time Frame: Assessments are collected within the first 72 hours following admission.
|
Evaluated and tested for the same as the primary, secondary and first exploratory outcomes (i.e., circulatory, respiratory, and renal failure risks) using the respective assessments by treating physicians and treating nurses with paired time-point assessments.
|
Assessments are collected within the first 72 hours following admission.
|
|
Comparison of predictive accuracy of treating versus non-treating physicians
Time Frame: Assessments are collected within the first 72 hours following admission.
|
Evaluated and tested for the same as the primary, secondary and first exploratory outcomes (i.e., circulatory, respiratory, and renal failure risks) using the respective assessments by clinicians (treating) and solely relying on EHR data (non-treating) with paired time-point assessments.
|
Assessments are collected within the first 72 hours following admission.
|
|
Calibration analysis of clinician prediction scores
Time Frame: Assessments are collected within the first 72 hours following admission.
|
Calibration analysis of clinician prediction scores.
The study assesses the prediction capabilities of a collective of clinicians.
However, humans might be ill-calibrated amongst each other with respect to providing probability estimates.
Assessing the collective's performance without calibration of the individuals amongst each other might underestimate the actual predictive capabilities of the clinicians if they were well-calibrated.
We propose to re-assess the primary and secondary outcomes but additionally perform a risk score calibration amongst physicians.
|
Assessments are collected within the first 72 hours following admission.
|
|
Exploratory analysis of patterns (e.g., systematic errors/biases) of predictive performance in human and machine assessors
Time Frame: Assessments are collected within the first 72 hours following admission.
|
|
Assessments are collected within the first 72 hours following admission.
|
Collaborators and Investigators
This is where you will find people and organizations involved with this study.
Sponsor
Collaborators
Investigators
- Principal Investigator: Martin Faltys, Dr. med., Insel Gruppe AG, University Hospital Bern
Study record dates
These dates track the progress of study record and summary results submissions to ClinicalTrials.gov. Study records and reported results are reviewed by the National Library of Medicine (NLM) to make sure they meet specific quality control standards before being posted on the public website.
Study Major Dates
Study Start (Actual)
November 18, 2024
Primary Completion (Actual)
May 15, 2025
Study Completion (Estimated)
May 15, 2026
Study Registration Dates
First Submitted
November 18, 2024
First Submitted That Met QC Criteria
February 3, 2025
First Posted (Actual)
February 7, 2025
Study Record Updates
Last Update Posted (Estimated)
August 28, 2025
Last Update Submitted That Met QC Criteria
August 27, 2025
Last Verified
November 1, 2024
More Information
Terms related to this study
Additional Relevant MeSH Terms
- Urogenital Diseases
- Pathologic Processes
- Male Urogenital Diseases
- Kidney Diseases
- Urologic Diseases
- Female Urogenital Diseases
- Female Urogenital Diseases and Pregnancy Complications
- Respiratory Tract Diseases
- Respiration Disorders
- Pathological Conditions, Signs and Symptoms
- Respiratory Insufficiency
- Renal Insufficiency
- Shock
Other Study ID Numbers
- BASEC-01046
Plan for Individual participant data (IPD)
Plan to Share Individual Participant Data (IPD)?
UNDECIDED
Drug and device information, study documents
Studies a U.S. FDA-regulated drug product
No
Studies a U.S. FDA-regulated device product
No
product manufactured in and exported from the U.S.
No
This information was retrieved directly from the website clinicaltrials.gov without any changes. If you have any requests to change, remove or update your study details, please contact register@clinicaltrials.gov. As soon as a change is implemented on clinicaltrials.gov, this will be updated automatically on our website as well.