此页面是自动翻译的,不保证翻译的准确性。请参阅 英文版 对于源文本。

Large Language Models Versus Anesthesiologists for ASA Physical Status Classification (ASA-LLM)

2026年7月6日 更新者:dilara gocmen、Marmara University Pendik Training and Research Hospital

Comparison of Clinical Assessment and Large Language Models in Preoperative Risk Classification: A Retrospective Analysis of ChatGPT, DeepSeek, Gemini, and Claude in ASA Physical Status Classification

The American Society of Anesthesiologists Physical Status (ASA-PS) classification is a cornerstone of preoperative risk assessment, yet interrater variability among clinicians is well documented. Large language models (LLMs) have recently demonstrated expert-level performance in several clinical classification tasks, including ASA-PS assignment.

This retrospective observational study evaluates whether four widely used LLMs - ChatGPT, DeepSeek, Gemini, and Claude - can accurately and consistently assign ASA-PS classes from structured, fully anonymized clinical vignettes derived from real preoperative anesthesia evaluations, using a consensus of senior anesthesiologists as the reference standard.

No patient data will be transmitted to third-party platforms. Clinical information will be converted by the investigators into de-identified structured vignettes containing only age range, sex, body mass index range, presence or absence of systemic diseases, functional capacity, and the major/minor nature of the planned surgery, in full compliance with national data protection legislation (KVKK).

研究概览

地位

尚未招聘

详细说明

Adult patients who underwent preoperative anesthesia evaluation before elective surgery at Marmara University Pendik Training and Research Hospital will be included retrospectively. For each patient, demographic data (age, sex, body mass index), systemic comorbidities (hypertension, diabetes mellitus, coronary artery disease, chronic obstructive pulmonary disease, and others), functional capacity (metabolic equivalents, MET), type of planned surgery (major/minor), and the ASA-PS class assigned by the attending anesthesiologist will be recorded.

Clinical data will be anonymized and converted into structured clinical vignettes by the investigators. Vignettes will contain no identifiers, dates, protocol numbers, or rare diagnostic combinations that could directly or indirectly identify a patient.

Standardization of the LLM assessment process: To ensure independence between assessments, each vignette will be evaluated in a separate, history-free session. A new conversation will be initiated in the relevant model for every patient vignette, thereby eliminating the possibility that the model is influenced by its responses to previous vignettes (context anchoring). The ASA-PS class assigned to one vignette will not be carried over as context into the evaluation of any subsequent vignette. Each vignette will be presented to all four models using an identical, standardized prompt requesting only an ASA-PS class (I-VI) with a brief rationale, in a strictly defined output format. Model outputs will play no role in clinical decision-making. External information retrieval by the models will be disabled, and all queries will be completed within a narrow time window to minimize variability in model versions.

Each vignette will be submitted to each model once (single querying). Consequently, the intra-model test-retest reliability of the LLMs will not be assessed; this is acknowledged as a study limitation, consistent with the probabilistic nature of large language models, which may produce between-session variability in their outputs.

Model versions: The current version of each model available at the time of data collection will be used - ChatGPT (GPT-5.5, OpenAI), Gemini (Gemini 3.5, Google DeepMind), DeepSeek (DeepSeek V4, DeepSeek AI), and Claude (Claude Opus 4.8, Anthropic). These versions reflect the versions current at the time of protocol submission; the most recent stable version of each model accessible during data collection will be used, and the exact version and access date will be recorded. Because publicly available chat interfaces may perform automatic background routing to different model tiers, this is acknowledged as a reproducibility limitation.

The reference standard ASA-PS class will be determined by an independent, blinded panel of at least three senior anesthesiologists; consensus or majority vote will define the reference classification.

Statistical analysis: The primary (confirmatory) analysis will quantify the agreement between each LLM and the reference standard using quadratic weighted Cohen's kappa, respecting the ordinal structure of ASA-PS. Multi-rater agreement across the four models and the human raters will be assessed with Fleiss' kappa. Pairwise accuracy comparisons among the four models (six pairwise contrasts) will be treated as secondary/exploratory analyses and compared with McNemar or permutation tests for paired data, applying correction for multiple comparisons (e.g., Bonferroni or Holm); 95% confidence intervals will be estimated by bootstrap methods. Prespecified subgroup analyses include ASA III-IV boundary cases, multimorbidity burden, major versus minor surgery, and rater experience.

Primary hypothesis: The ASA-PS assignments of the LLMs (ChatGPT, DeepSeek, Gemini, and Claude) will show at least good agreement with the reference standard (weighted kappa ≥ 0.60). Secondary hypothesis: LLM errors will cluster in specific subgroups (e.g., the ASA III-IV boundary, multimorbid patients).

研究类型

观察性的

注册 (估计的)

350

联系人和位置

本节提供了进行研究的人员的详细联系信息,以及有关进行该研究的地点的信息。

学习联系方式

参与标准

研究人员寻找符合特定描述的人,称为资格标准。这些标准的一些例子是一个人的一般健康状况或先前的治疗。

资格标准

适合学习的年龄

  • 成人
  • 年长者

接受健康志愿者

不

取样方法

概率样本

研究人群

Adult patients who underwent preoperative anesthesia evaluation before elective surgery at a tertiary university hospital in Istanbul, Turkey.

描述

Inclusion Criteria:

  • Age 18 years or older
  • Planned elective surgery
  • Completed preoperative anesthesia evaluation

Exclusion Criteria:

  • Emergency surgical procedures
  • ASA VI (brain death)
  • Incomplete clinical records

学习计划

本节提供研究计划的详细信息,包括研究的设计方式和研究的衡量标准。

研究是如何设计的?

设计细节

队列和干预

团体/队列
elective surgery patients
Adult patients (≥18 years) who underwent preoperative anesthesia evaluation before elective surgery. Anonymized structured vignettes derived from their records will be classified by four LLMs (ChatGPT, DeepSeek, Gemini, Claude) and by a blinded senior anesthesiologist panel serving as the reference standard.

研究衡量的是什么?

主要结果指标

结果测量
措施说明
大体时间
Agreement between LLM-assigned and reference-standard ASA-PS class
大体时间:Through study completion, an average of 3 months
Quadratic weighted Cohen's kappa between each large language model's ASA-PS assignment (ChatGPT, DeepSeek, Gemini, Claude) and the reference standard defined by consensus of a blinded panel of at least three senior anesthesiologists. Agreement of at least "good" level (weighted kappa ≥ 0.60) is hypothesized.
Through study completion, an average of 3 months

次要结果测量

结果测量
措施说明
大体时间
Overall classification accuracy of each LLM
大体时间:Through study completion, an average of 3 months
Proportion of vignettes in which the LLM-assigned ASA-PS class exactly matches the reference standard, with exploratory pairwise comparisons among the four models (McNemar/permutation tests, corrected for multiple comparisons)
Through study completion, an average of 3 months
Subgroup error patterns
大体时间:Through study completion, an average of 3 months
Frequency and direction (over- vs. under-classification) of LLM misclassifications in prespecified subgroups: ASA III-IV boundary, multimorbidity, major vs. minor surgery
Through study completion, an average of 3 months

合作者和调查者

在这里您可以找到参与这项研究的人员和组织。

出版物和有用的链接

负责输入研究信息的人员自愿提供这些出版物。这些可能与研究有关。

研究记录日期

这些日期跟踪向 ClinicalTrials.gov 提交研究记录和摘要结果的进度。研究记录和报告的结果由国家医学图书馆 (NLM) 审查,以确保它们在发布到公共网站之前符合特定的质量控制标准。

研究主要日期

学习开始 (估计的)

2026年7月21日

初级完成 (估计的)

2026年8月21日

研究完成 (估计的)

2026年10月21日

研究注册日期

首次提交

2026年7月6日

首先提交符合 QC 标准的

2026年7月6日

首次发布 (实际的)

2026年7月10日

研究记录更新

最后更新发布 (实际的)

2026年7月10日

上次提交的符合 QC 标准的更新

2026年7月6日

最后验证

2026年7月1日

更多信息

与本研究相关的术语

计划个人参与者数据 (IPD)

计划共享个人参与者数据 (IPD)?

不

药物和器械信息、研究文件

研究美国 FDA 监管的药品

不

研究美国 FDA 监管的设备产品

不

此信息直接从 clinicaltrials.gov 网站检索,没有任何更改。如果您有任何更改、删除或更新研究详细信息的请求,请联系 register@clinicaltrials.gov. clinicaltrials.gov 上实施更改,我们的网站上也会自动更新.

订阅