Questa pagina è stata tradotta automaticamente e l'accuratezza della traduzione non è garantita. Si prega di fare riferimento al Versione inglese per un testo di partenza.

Large Language Models Versus Anesthesiologists for ASA Physical Status Classification (ASA-LLM)

6 luglio 2026 aggiornato da: dilara gocmen, Marmara University Pendik Training and Research Hospital

Comparison of Clinical Assessment and Large Language Models in Preoperative Risk Classification: A Retrospective Analysis of ChatGPT, DeepSeek, Gemini, and Claude in ASA Physical Status Classification

The American Society of Anesthesiologists Physical Status (ASA-PS) classification is a cornerstone of preoperative risk assessment, yet interrater variability among clinicians is well documented. Large language models (LLMs) have recently demonstrated expert-level performance in several clinical classification tasks, including ASA-PS assignment.

This retrospective observational study evaluates whether four widely used LLMs - ChatGPT, DeepSeek, Gemini, and Claude - can accurately and consistently assign ASA-PS classes from structured, fully anonymized clinical vignettes derived from real preoperative anesthesia evaluations, using a consensus of senior anesthesiologists as the reference standard.

No patient data will be transmitted to third-party platforms. Clinical information will be converted by the investigators into de-identified structured vignettes containing only age range, sex, body mass index range, presence or absence of systemic diseases, functional capacity, and the major/minor nature of the planned surgery, in full compliance with national data protection legislation (KVKK).

Panoramica dello studio

Descrizione dettagliata

Adult patients who underwent preoperative anesthesia evaluation before elective surgery at Marmara University Pendik Training and Research Hospital will be included retrospectively. For each patient, demographic data (age, sex, body mass index), systemic comorbidities (hypertension, diabetes mellitus, coronary artery disease, chronic obstructive pulmonary disease, and others), functional capacity (metabolic equivalents, MET), type of planned surgery (major/minor), and the ASA-PS class assigned by the attending anesthesiologist will be recorded.

Clinical data will be anonymized and converted into structured clinical vignettes by the investigators. Vignettes will contain no identifiers, dates, protocol numbers, or rare diagnostic combinations that could directly or indirectly identify a patient.

Standardization of the LLM assessment process: To ensure independence between assessments, each vignette will be evaluated in a separate, history-free session. A new conversation will be initiated in the relevant model for every patient vignette, thereby eliminating the possibility that the model is influenced by its responses to previous vignettes (context anchoring). The ASA-PS class assigned to one vignette will not be carried over as context into the evaluation of any subsequent vignette. Each vignette will be presented to all four models using an identical, standardized prompt requesting only an ASA-PS class (I-VI) with a brief rationale, in a strictly defined output format. Model outputs will play no role in clinical decision-making. External information retrieval by the models will be disabled, and all queries will be completed within a narrow time window to minimize variability in model versions.

Each vignette will be submitted to each model once (single querying). Consequently, the intra-model test-retest reliability of the LLMs will not be assessed; this is acknowledged as a study limitation, consistent with the probabilistic nature of large language models, which may produce between-session variability in their outputs.

Model versions: The current version of each model available at the time of data collection will be used - ChatGPT (GPT-5.5, OpenAI), Gemini (Gemini 3.5, Google DeepMind), DeepSeek (DeepSeek V4, DeepSeek AI), and Claude (Claude Opus 4.8, Anthropic). These versions reflect the versions current at the time of protocol submission; the most recent stable version of each model accessible during data collection will be used, and the exact version and access date will be recorded. Because publicly available chat interfaces may perform automatic background routing to different model tiers, this is acknowledged as a reproducibility limitation.

The reference standard ASA-PS class will be determined by an independent, blinded panel of at least three senior anesthesiologists; consensus or majority vote will define the reference classification.

Statistical analysis: The primary (confirmatory) analysis will quantify the agreement between each LLM and the reference standard using quadratic weighted Cohen's kappa, respecting the ordinal structure of ASA-PS. Multi-rater agreement across the four models and the human raters will be assessed with Fleiss' kappa. Pairwise accuracy comparisons among the four models (six pairwise contrasts) will be treated as secondary/exploratory analyses and compared with McNemar or permutation tests for paired data, applying correction for multiple comparisons (e.g., Bonferroni or Holm); 95% confidence intervals will be estimated by bootstrap methods. Prespecified subgroup analyses include ASA III-IV boundary cases, multimorbidity burden, major versus minor surgery, and rater experience.

Primary hypothesis: The ASA-PS assignments of the LLMs (ChatGPT, DeepSeek, Gemini, and Claude) will show at least good agreement with the reference standard (weighted kappa ≥ 0.60). Secondary hypothesis: LLM errors will cluster in specific subgroups (e.g., the ASA III-IV boundary, multimorbid patients).

Tipo di studio

Osservativo

Iscrizione (Stimato)

350

Contatti e Sedi

Questa sezione fornisce i recapiti di coloro che conducono lo studio e informazioni su dove viene condotto lo studio.

Contatto studio

Criteri di partecipazione

I ricercatori cercano persone che corrispondano a una certa descrizione, chiamata criteri di ammissibilità. Alcuni esempi di questi criteri sono le condizioni generali di salute di una persona o trattamenti precedenti.

Criteri di ammissibilità

Età idonea allo studio

  • Adulto
  • Adulto più anziano

Accetta volontari sani

No

Metodo di campionamento

Campione di probabilità

Popolazione di studio

Adult patients who underwent preoperative anesthesia evaluation before elective surgery at a tertiary university hospital in Istanbul, Turkey.

Descrizione

Inclusion Criteria:

  • Age 18 years or older
  • Planned elective surgery
  • Completed preoperative anesthesia evaluation

Exclusion Criteria:

  • Emergency surgical procedures
  • ASA VI (brain death)
  • Incomplete clinical records

Piano di studio

Questa sezione fornisce i dettagli del piano di studio, compreso il modo in cui lo studio è progettato e ciò che lo studio sta misurando.

Come è strutturato lo studio?

Dettagli di progettazione

Coorti e interventi

Gruppo / Coorte
elective surgery patients
Adult patients (≥18 years) who underwent preoperative anesthesia evaluation before elective surgery. Anonymized structured vignettes derived from their records will be classified by four LLMs (ChatGPT, DeepSeek, Gemini, Claude) and by a blinded senior anesthesiologist panel serving as the reference standard.

Cosa sta misurando lo studio?

Misure di risultato primarie

Misura del risultato
Misura Descrizione
Lasso di tempo
Agreement between LLM-assigned and reference-standard ASA-PS class
Lasso di tempo: Through study completion, an average of 3 months
Quadratic weighted Cohen's kappa between each large language model's ASA-PS assignment (ChatGPT, DeepSeek, Gemini, Claude) and the reference standard defined by consensus of a blinded panel of at least three senior anesthesiologists. Agreement of at least "good" level (weighted kappa ≥ 0.60) is hypothesized.
Through study completion, an average of 3 months

Misure di risultato secondarie

Misura del risultato
Misura Descrizione
Lasso di tempo
Overall classification accuracy of each LLM
Lasso di tempo: Through study completion, an average of 3 months
Proportion of vignettes in which the LLM-assigned ASA-PS class exactly matches the reference standard, with exploratory pairwise comparisons among the four models (McNemar/permutation tests, corrected for multiple comparisons)
Through study completion, an average of 3 months
Subgroup error patterns
Lasso di tempo: Through study completion, an average of 3 months
Frequency and direction (over- vs. under-classification) of LLM misclassifications in prespecified subgroups: ASA III-IV boundary, multimorbidity, major vs. minor surgery
Through study completion, an average of 3 months

Collaboratori e investigatori

Qui è dove troverai le persone e le organizzazioni coinvolte in questo studio.

Pubblicazioni e link utili

La persona responsabile dell'inserimento delle informazioni sullo studio fornisce volontariamente queste pubblicazioni. Questi possono riguardare qualsiasi cosa relativa allo studio.

Studiare le date dei record

Queste date tengono traccia dell'avanzamento della registrazione dello studio e dell'invio dei risultati di sintesi a ClinicalTrials.gov. I record degli studi e i risultati riportati vengono esaminati dalla National Library of Medicine (NLM) per assicurarsi che soddisfino specifici standard di controllo della qualità prima di essere pubblicati sul sito Web pubblico.

Studia le date principali

Inizio studio (Stimato)

21 luglio 2026

Completamento primario (Stimato)

21 agosto 2026

Completamento dello studio (Stimato)

21 ottobre 2026

Date di iscrizione allo studio

Primo inviato

6 luglio 2026

Primo inviato che soddisfa i criteri di controllo qualità

6 luglio 2026

Primo Inserito (Effettivo)

10 luglio 2026

Aggiornamenti dei record di studio

Ultimo aggiornamento pubblicato (Effettivo)

10 luglio 2026

Ultimo aggiornamento inviato che soddisfa i criteri QC

6 luglio 2026

Ultimo verificato

1 luglio 2026

Maggiori informazioni

Termini relativi a questo studio

Piano per i dati dei singoli partecipanti (IPD)

Hai intenzione di condividere i dati dei singoli partecipanti (IPD)?

NO

Informazioni su farmaci e dispositivi, documenti di studio

Studia un prodotto farmaceutico regolamentato dalla FDA degli Stati Uniti

No

Studia un dispositivo regolamentato dalla FDA degli Stati Uniti

No

Queste informazioni sono state recuperate direttamente dal sito web clinicaltrials.gov senza alcuna modifica. In caso di richieste di modifica, rimozione o aggiornamento dei dettagli dello studio, contattare register@clinicaltrials.gov. Non appena verrà implementata una modifica su clinicaltrials.gov, questa verrà aggiornata automaticamente anche sul nostro sito web .

Sottoscrivi