Diagnostic Accuracy of Two Large Language Models in Turkish Emergency Department Anamnesis Notes (LLM-ED-DX-TR)
Diagnostic Accuracy of Two Large Language Models Against a Blinded Specialist Consensus Standard in Turkish Emergency Department Notes: A Retrospective Study of 600 Cases
This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes.
The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at case closure, and to the subsequent clinical course. Cases without chapter-level majority agreement are excluded without replacement.
Both models are queried once per note with a single locked prompt at temperature 0 in stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. The primary outcome is the proportion of cases in which each model's rank-1 diagnosis matches the reference standard at ICD-10 chapter level, reported with a Wilson 95% confidence interval. Registered secondary outcome measures are chapter-level Cohen's kappa between each model's rank-1 diagnosis and the reference standard; top-3 chapter accuracy for each model; and chapter-level concordance between the closure ICD-10 code and the reference standard. Additional prespecified analyses set out in the statistical analysis plan (paired between-model difference, three-character accuracy, note-length association, confidence calibration and model-to-model agreement) are reported in the primary publication.
The ICD-10 code entered at case closure is characterised against the same reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed. The analysis plan was finalised and frozen before any accuracy computation. Reporting follows STARD-AI 2025.
調査の概要
状態
状態
条件
条件
詳細な説明
STUDY DESIGN: Retrospective diagnostic accuracy study, STARD-AI 2025 reporting, single centre, cohort design.
AI INDEX TESTS: (1) GPT-4.1 (model version gpt-4.1-2025-04-14; OpenAI API). (2) Claude Sonnet 4.6 (model version claude-sonnet-4-6; Anthropic API). Both accessed via the providers' developer application programming interfaces from Python. Temperature = 0. Zero-shot direct prompting with a single locked prompt version; stateless single-turn sessions with no cross-case context, no retrieval augmentation, no external tools and no extended-reasoning mode. No task-specific fine-tuning or additional training was applied; the models were used as released.
MODEL INTERPRETABILITY: Interpretability analyses such as SHAP, Grad-CAM or layer-attribution visualisation are not applicable to this study. Because GPT-4.1 and Claude Sonnet 4.6 are accessed as black-box models through proprietary, closed-source commercial interfaces, internal weights, gradients and attention structures are inaccessible for post-hoc interpretability computation.
REFERENCE STANDARD: Three board-certified emergency medicine specialists independently assess each anonymized note, blinded to one another, to the code entered by the treating physician, and to the subsequent clinical course. The primary diagnosis assigned by at least two of three assessors, reduced to ICD-10 chapter level, constitutes the reference standard. Cases in which all three assessors assign different chapters are excluded without replacement. No joint calibration session was held and no adjudication round was performed; each assessor coded once, according to their own clinical judgement.
DATA PRIVACY: All anamnesis notes are de-identified before processing; direct patient identifiers are removed and no patient name is present in any note at any stage. Each case carries a study-specific sequential number that is not a hospital record number, and no file linking study numbers to patient identities was created or retained. Note text is transmitted to commercial application programming interfaces operated by providers established outside Turkiye; all queries are issued through the providers' developer interfaces in stateless single-turn calls, and no patient identifier is present in any submitted text. De-identified notes are stored in an encrypted, access-restricted database. Conducted in accordance with Turkish Personal Data Protection Law no. 6698.
REPRODUCIBILITY OF THE INDEX TEST: The statistical analysis plan specified a test-retest assessment of within-model reproducibility. A random subset of 20 cases was drawn from the analysis set with a fixed seed recorded before the re-run, and both models were re-queried on those notes on 7 August 2026, after an interval of 3 days and 18 hours from the primary run (protocol minimum 48 hours), using the same prompt content, the same model identifiers and the same sampling parameters; the single-query-per-note statement above refers to the primary run. The prespecified measure is the proportion of cases in which the rank-1 code is identical between runs, at three-character and at ICD-10 chapter level. Two conditions differed from the primary run and are recorded in the deviation log: the byte-exact prompt file used in the primary run could not be recovered, only its SHA-256 digest having been retained, so the re-run used a prompt of identical content but unverified byte identity; and the structured-output mechanism for GPT-4.1 was JSON schema mode at re-run rather than the JSON object mode used originally, which the API rejected. The analysis is therefore reported as consistency of re-execution rather than strict prompt-identical reproducibility. This re-run is the last date of data collection and determines the study completion date.
STUDY DATES: The Actual Study Start Date (1 May 2026) denotes the beginning of the retrospective encounter window from which archived notes were drawn, not the start of data collection. Ethics approval (Clinical Research Ethics Committee of Marmara University, protocol 09.2026.26-0514) was granted on 14 May 2026. Because the study is retrospective, every note analysed was already present in the hospital record system when it was retrieved; no data were generated prospectively and no patient was enrolled.
STUDY FLOW: 630 consecutive eligible notes were screened and coded by all three assessors. Ten notes (cases 621-630) fell beyond the ethics-approved ceiling of 600 analysable cases and were excluded before analysis, leaving an assessment window of 620 on which inter-assessor agreement is reported. Within that window 20 notes had no chapter-level majority among the three assessors and were excluded without replacement, giving a primary analysis set of 600. A further 4 notes had no majority three-character code, giving 596 for the secondary three-character analysis.
STATISTICAL ANALYSIS: Analyses are performed in Python 3.11 (pandas, statsmodels, scipy) following a statistical analysis plan finalised and frozen before any accuracy computation; selected estimates are independently recomputed in jamovi by a second investigator using a prespecified verification checklist.
PATIENT AND PUBLIC INVOLVEMENT: Not applicable. This retrospective study uses existing anonymized records; there was no patient or public involvement in design or conduct.
DATA SHARING: De-identified data are available from the principal investigator on reasonable request, subject to institutional approval. Eight supplementary files are provided with the primary publication: the statistical analysis plan with its deviation log; the full prompt text with its recorded SHA-256 digest; the data-preparation and analysis code; the data dictionary; the completed STARD-AI 2025 reporting checklist; the jamovi verification checklist used for independent recomputation of selected estimates; the technical specification of the two index tests; and the full chapter-level confusion matrices for both models.
研究の種類
研究の種類
入学 (実際)
入学
連絡先と場所
研究連絡先
研究連絡先
- 名前:Emir Ünal, Assistant Professor
- 電話番号:+905327766010
- メール:emirunal@gmail.com
研究連絡先のバックアップ
- 名前:Emir Unal, Assistant Professor
- メール:emirunal@gmail.com
研究場所
-
-
Istanbul
-
Istanbul、Istanbul、トルコ(Türkiye)、34899
- Marmara University Pendik Training and Research Hospital
-
-
参加基準
適格基準
適格基準
就学可能な年齢
- 大人
- 高齢者
健康ボランティアの受け入れ
サンプリング方法
調査対象母集団
説明
INCLUSION CRITERIA:
Adult patients (aged 18 years and older) presenting to the emergency department, evaluated in the ambulatory (green/yellow triage) area.
A free-text electronic anamnesis note entered at presentation in the hospital information system (HBYS). No minimum note length and no "sufficient information for diagnosis" requirement was applied, because such a criterion preferentially retains more readily classifiable cases; note length was treated as a covariate rather than as an eligibility threshold. A note was excluded only if all three of the following were absent: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
An ICD-10 code entered by the treating emergency physician at case closure. Cases in which this entry was absent or did not form a valid ICD-10 code were retained in the analysis set and counted in the denominator of the closure-code analyses.
EXCLUSION CRITERIA:
Notes lacking all three of the following: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
Pediatric cases (age under 18 years).
Patients critically ill and triaged to high-acuity resuscitation areas (Emergency Severity Index [ESI] level 1).
Clinical notes containing residual identifying information that cannot be fully de-identified, preventing compliance with data privacy regulations.
Non-independent clinical notes consisting solely of a brief cross-reference to a prior hospital visit without a new history entry.
研究計画
研究はどのように設計されていますか?
デザインの詳細
グループ/コホートの数
コホートと介入
グループ/コホートグループ/コホート |
|---|
|
Emergency Department Patient Cohort
Consecutive adult patients (aged 18 years and older) evaluated in the ambulatory (green/yellow triage) area of the emergency department, who had a free-text electronic anamnesis note recorded at presentation and an ICD-10 code entered by the treating physician at case closure.
No note-completeness or minimum-length requirement was applied.
The closure code is characterised against the reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed.
|
この研究は何を測定していますか?
主要な結果の測定
主要な結果の測定
結果測定 |
メジャーの説明 |
時間枠 |
|---|---|---|
|
Diagnostic Accuracy of GPT-4.1 for ICD-10 Chapter-Level Diagnosis
時間枠:At the single index-test run on 3 August 2026
|
Proportion of cases in which the GPT-4.1 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories).
Range: 0 to 1.00.
|
At the single index-test run on 3 August 2026
|
|
Diagnostic Accuracy of Claude Sonnet 4.6 for ICD-10 Chapter-Level Diagnosis
時間枠:At the single index-test run on 3 August 2026
|
Proportion of cases in which the Claude Sonnet 4.6 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories).
Range: 0 to 1.00.
|
At the single index-test run on 3 August 2026
|
二次結果の測定
二次結果の測定
結果測定 |
メジャーの説明 |
時間枠 |
|---|---|---|
|
Cohen's Kappa Between GPT-4.1 Primary Diagnosis and the Reference Standard
時間枠:At the single index-test run on 3 August 2026
|
Kappa coefficient measuring agreement between the GPT-4.1 rank-1 ICD-10 chapter and the 3-specialist reference standard.
Interpreted per Landis & Koch (1977): <=0.20 slight; 0.21-0.40
fair; 0.41-0.60
moderate; 0.61-0.80
substantial; >0.80 almost perfect.
Range: -1.00 to 1.00.
|
At the single index-test run on 3 August 2026
|
|
Cohen's Kappa Between Claude Sonnet 4.6 Primary Diagnosis and the Reference Standard
時間枠:At the single index-test run on 3 August 2026
|
Kappa coefficient measuring agreement between the Claude Sonnet 4.6 rank-1 ICD-10 chapter and the 3-specialist reference standard.
Interpreted per Landis & Koch (1977): <=0.20 slight; 0.21-0.40
fair; 0.41-0.60
moderate; 0.61-0.80
substantial; >0.80 almost perfect.
Range: -1.00 to 1.00.
|
At the single index-test run on 3 August 2026
|
|
Top-3 Diagnostic Accuracy of GPT-4.1
時間枠:At the single index-test run on 3 August 2026
|
Proportion of cases in which the ICD-10 chapter of the reference standard diagnosis appears anywhere within the ranked list of three differential diagnoses returned by GPT-4.1.
Range: 0 to 1.00.
Cases in which no valid closure code was entered (5 of 600) are retained in the denominator; the figure restricted to resolvable entries is reported alongside.
|
At the single index-test run on 3 August 2026
|
|
Top-3 Diagnostic Accuracy of Claude Sonnet 4.6
時間枠:At the single index-test run on 3 August 2026
|
Proportion of cases in which the ICD-10 chapter of the reference standard diagnosis appears anywhere within the ranked list of three differential diagnoses returned by Claude Sonnet 4.6.
Range: 0 to 1.00.
|
At the single index-test run on 3 August 2026
|
|
Chapter-Level Concordance Between the Closure ICD-10 Code and the Reference Standard
時間枠:At the original clinical encounter (retrospective data spanning 1 May to 3 August 2026)
|
Proportion of cases in which the ICD-10 code entered by the treating emergency physician at case closure matches the 3-specialist reference standard at the chapter level.
This is reported as a descriptive benchmark of routine coding practice and is not a comparator: the closure code was entered after investigation, whereas the reference standard was constructed from the presentation note alone, to which the assessors were restricted.
Range: 0 to 1.00.
|
At the original clinical encounter (retrospective data spanning 1 May to 3 August 2026)
|
協力者と研究者
捜査官
捜査官
- 主任研究者:Emir Ünal、Marmara University
出版物と役立つリンク
一般刊行物
- Kanjee Z, Crowe B, Rodman A. Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA. 2023 Jul 3;330(1):78-80. doi: 10.1001/jama.2023.8288.
- Sounderajah V, Guni A, Liu X, Collins GS, Karthikesalingam A, Markar SR, Golub RM, Denniston AK, Shetty S, Moher D, Bossuyt PM, Darzi A, Ashrafian H; STARD-AI Steering Committee. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. 2025 Oct;31(10):3283-3289. doi: 10.1038/s41591-025-03953-8. Epub 2025 Sep 15.
- Newman-Toker DE, Peterson SM, Badihian S, Hassoon A, Nassery N, Parizadeh D, Wilson LM, Jia Y, Omron R, Tharmarajah S, Guerin L, Bastani PB, Fracica EA, Kotwal S, Robinson KA. Diagnostic Errors in the Emergency Department: A Systematic Review [Internet]. Rockville (MD): Agency for Healthcare Research and Quality (US); 2022 Dec. Report No.: 22(23)-EHC043. Available from http://www.ncbi.nlm.nih.gov/books/NBK588118/
- Wei J et al. Chain-of-thought prompting elicits reasoning in LLMs. NeurIPS. 2022;35:24824-24837.
- Niset A, Melot I, Pireau M, Englebert A, Scius N, Flament J, El Hadwe S, Al Barajraji M, Thonon H, Barrit S. Grounded large language models for diagnostic prediction in real-world emergency department settings. JAMIA Open. 2025 Oct 21;8(5):ooaf119. doi: 10.1093/jamiaopen/ooaf119. eCollection 2025 Oct.
- Williams CYK, Miao BY, Kornblith AE, Butte AJ. Evaluating the use of large language models to provide clinical recommendations in the Emergency Department. Nat Commun. 2024 Oct 8;15(1):8236. doi: 10.1038/s41467-024-52415-1.
- Hoppe JM, Auer MK, Struven A, Massberg S, Stremmel C. ChatGPT With GPT-4 Outperforms Emergency Department Physicians in Diagnostic Accuracy: Retrospective Analysis. J Med Internet Res. 2024 Jul 8;26:e56110. doi: 10.2196/56110.
- Takita H, Kabata D, Walston SL, Tatekawa H, Saito K, Tsujimoto Y, Miki Y, Ueda D. A systematic review and meta-analysis of diagnostic performance comparison between generative AI and physicians. NPJ Digit Med. 2025 Mar 22;8(1):175. doi: 10.1038/s41746-025-01543-z.
- Shan G, Chen X, Wang C, Liu L, Gu Y, Jiang H, Shi T. Comparing Diagnostic Accuracy of Clinical Professionals and Large Language Models: Systematic Review and Meta-Analysis. JMIR Med Inform. 2025 Apr 25;13:e64963. doi: 10.2196/64963.
- Taylor RA, Sangal RB, Smith ME, Haimovich AD, Rodman A, Iscoe MS, Pavuluri SK, Rose C, Janke AT, Wright DS, Socrates V, Declan A. Leveraging artificial intelligence to reduce diagnostic errors in emergency medicine: Challenges, opportunities, and future directions. Acad Emerg Med. 2025 Mar;32(3):327-339. doi: 10.1111/acem.15066. Epub 2024 Dec 15.
研究記録日
主要日程の研究
研究開始 (実際)
研究開始
一次修了 (実際)
一次修了
研究の完了 (実際)
研究の完了
試験登録日
最初に提出
最初に提出
QC基準を満たした最初の提出物
QC基準を満たした最初の提出物
最初の投稿 (実際)
最初の投稿
学習記録の更新
投稿された最後の更新 (実際)
投稿された最後の更新
QC基準を満たした最後の更新が送信されました
QC基準を満たした最後の更新が送信されました
最終確認日
最終確認日
詳しくは
本研究に関する用語
追加の関連 MeSH 用語
その他の研究ID番号
その他の研究ID番号
- 09.2026.26-0514
医薬品およびデバイス情報、研究文書
米国FDA規制医薬品の研究
米国FDA規制機器製品の研究
この情報は、Web サイト clinicaltrials.gov から変更なしで直接取得したものです。研究の詳細を変更、削除、または更新するリクエストがある場合は、register@clinicaltrials.gov。 までご連絡ください。 clinicaltrials.gov に変更が加えられるとすぐに、ウェブサイトでも自動的に更新されます。