このページは自動翻訳されたものであり、翻訳の正確性は保証されていません。を参照してください。 英語版 ソーステキスト用。

Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models

2026年7月23日 更新者:Mariam Ahmed Hossam、Cairo University

Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models in Endodontics Using Expert Consensus as the Reference Standard

This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.

調査の概要

詳細な説明

This prospective, blinded diagnostic accuracy study aims to compare the diagnostic accuracy and endodontic case difficulty assessment of three LLMs-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-with expert consensus as the reference standard. Adult patients requiring primary endodontic treatment or retreatment will undergo routine clinical examination, pulp sensibility testing, and periapical radiographic assessment. Standardized clinical information and radiographs will be provided to each LLM and to four experienced endodontists independently. Expert consensus, defined as agreement among at least three of four experts, will serve as the reference standard. The primary outcomes are diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus for pulpal and periapical diagnosis and endodontic case difficulty assessment according to the American Association of Endodontists (AAE) criteria. The findings will provide evidence regarding the reliability of LLMs as clinical decision-support tools in endodontics and their potential role in improving diagnostic consistency and treatment planning.

研究の種類

観察的

入学 (推定)

342

連絡先と場所

このセクションには、調査を実施する担当者の連絡先の詳細と、この調査が実施されている場所に関する情報が記載されています。

研究連絡先

参加基準

研究者は、適格基準と呼ばれる特定の説明に適合する人を探します。これらの基準のいくつかの例は、人の一般的な健康状態または以前の治療です。

適格基準

就学可能な年齢

  • 子
  • 大人
  • 高齢者

健康ボランティアの受け入れ

はい

サンプリング方法

確率サンプル

調査対象母集団

Systemically healthy adult patients (≥16 years) requiring primary endodontic treatment or retreatment at the Endodontic Department, Faculty of Dentistry, Cairo University, who provide informed consent.

説明

Inclusion Criteria:

  1. Age above 16 years old.
  2. Systemically healthy patient (ASA I or II).
  3. Requiring endodontic treatment or retreatment
  4. Patient's acceptance to participate in the study

Exclusion Criteria:

  1. Medically compromised patients.
  2. Pregnant women.
  3. Traumatic dental injuries
  4. Low quality periapical radiograph

研究計画

このセクションでは、研究がどのように設計され、研究が何を測定しているかなど、研究計画の詳細を提供します。

研究はどのように設計されていますか?

デザインの詳細

この研究は何を測定していますか?

主要な結果の測定

結果測定
メジャーの説明
時間枠
Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosis
時間枠:At baseline
Diagnostic performance of each large language model compared with the expert consensus reference standard. Accuracy, sensitivity, specificity, and agreement (Cohen's kappa) will be calculated according to the American Association of Endodontists (AAE) diagnostic criteria
At baseline

二次結果の測定

結果測定
メジャーの説明
時間枠
Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessment
時間枠:At baseline
Agreement between each large language model and the expert consensus in classifying endodontic case difficulty according to the AAE Endodontic Case Difficulty Assessment Guidelines. Overall accuracy and weighted Cohen's kappa will be calculated.
At baseline

協力者と研究者

ここでは、この調査に関係する人々や組織を見つけることができます。

スポンサー

研究記録日

これらの日付は、ClinicalTrials.gov への研究記録と要約結果の提出の進捗状況を追跡します。研究記録と報告された結果は、国立医学図書館 (NLM) によって審査され、公開 Web サイトに掲載される前に、特定の品質管理基準を満たしていることが確認されます。

主要日程の研究

研究開始 (推定)

2026年9月1日

一次修了 (推定)

2026年12月1日

研究の完了 (推定)

2027年1月1日

試験登録日

最初に提出

2026年7月10日

QC基準を満たした最初の提出物

2026年7月23日

最初の投稿 (実際)

2026年7月29日

学習記録の更新

投稿された最後の更新 (実際)

2026年7月29日

QC基準を満たした最後の更新が送信されました

2026年7月23日

最終確認日

2026年7月1日

詳しくは

本研究に関する用語

その他の研究ID番号

  • New ENDO7.1.1

個々の参加者データ (IPD) の計画

個々の参加者データ (IPD) を共有する予定はありますか?

未定

医薬品およびデバイス情報、研究文書

米国FDA規制医薬品の研究

いいえ

米国FDA規制機器製品の研究

いいえ

この情報は、Web サイト clinicaltrials.gov から変更なしで直接取得したものです。研究の詳細を変更、削除、または更新するリクエストがある場合は、register@clinicaltrials.gov。 までご連絡ください。 clinicaltrials.gov に変更が加えられるとすぐに、ウェブサイトでも自動的に更新されます。

購読する