Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models

July 23, 2026 updated by: Mariam Ahmed Hossam, Cairo University

Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models in Endodontics Using Expert Consensus as the Reference Standard

This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.

Study Overview

Detailed Description

This prospective, blinded diagnostic accuracy study aims to compare the diagnostic accuracy and endodontic case difficulty assessment of three LLMs-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-with expert consensus as the reference standard. Adult patients requiring primary endodontic treatment or retreatment will undergo routine clinical examination, pulp sensibility testing, and periapical radiographic assessment. Standardized clinical information and radiographs will be provided to each LLM and to four experienced endodontists independently. Expert consensus, defined as agreement among at least three of four experts, will serve as the reference standard. The primary outcomes are diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus for pulpal and periapical diagnosis and endodontic case difficulty assessment according to the American Association of Endodontists (AAE) criteria. The findings will provide evidence regarding the reliability of LLMs as clinical decision-support tools in endodontics and their potential role in improving diagnostic consistency and treatment planning.

Study Type

Observational

Enrollment (Estimated)

342

Contacts and Locations

This section provides the contact details for those conducting the study, and information on where this study is being conducted.

Study Contact

Participation Criteria

Researchers look for people who fit a certain description, called eligibility criteria. Some examples of these criteria are a person's general health condition or prior treatments.

Eligibility Criteria

Ages Eligible for Study

  • Child
  • Adult
  • Older Adult

Accepts Healthy Volunteers

Yes

Sampling Method

Probability Sample

Study Population

Systemically healthy adult patients (≥16 years) requiring primary endodontic treatment or retreatment at the Endodontic Department, Faculty of Dentistry, Cairo University, who provide informed consent.

Description

Inclusion Criteria:

  1. Age above 16 years old.
  2. Systemically healthy patient (ASA I or II).
  3. Requiring endodontic treatment or retreatment
  4. Patient's acceptance to participate in the study

Exclusion Criteria:

  1. Medically compromised patients.
  2. Pregnant women.
  3. Traumatic dental injuries
  4. Low quality periapical radiograph

Study Plan

This section provides details of the study plan, including how the study is designed and what the study is measuring.

How is the study designed?

Design Details

What is the study measuring?

Primary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosis
Time Frame: At baseline
Diagnostic performance of each large language model compared with the expert consensus reference standard. Accuracy, sensitivity, specificity, and agreement (Cohen's kappa) will be calculated according to the American Association of Endodontists (AAE) diagnostic criteria
At baseline

Secondary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessment
Time Frame: At baseline
Agreement between each large language model and the expert consensus in classifying endodontic case difficulty according to the AAE Endodontic Case Difficulty Assessment Guidelines. Overall accuracy and weighted Cohen's kappa will be calculated.
At baseline

Collaborators and Investigators

This is where you will find people and organizations involved with this study.

Study record dates

These dates track the progress of study record and summary results submissions to ClinicalTrials.gov. Study records and reported results are reviewed by the National Library of Medicine (NLM) to make sure they meet specific quality control standards before being posted on the public website.

Study Major Dates

Study Start (Estimated)

September 1, 2026

Primary Completion (Estimated)

December 1, 2026

Study Completion (Estimated)

January 1, 2027

Study Registration Dates

First Submitted

July 10, 2026

First Submitted That Met QC Criteria

July 23, 2026

First Posted (Actual)

July 29, 2026

Study Record Updates

Last Update Posted (Actual)

July 29, 2026

Last Update Submitted That Met QC Criteria

July 23, 2026

Last Verified

July 1, 2026

More Information

Terms related to this study

Other Study ID Numbers

  • New ENDO7.1.1

Plan for Individual participant data (IPD)

Plan to Share Individual Participant Data (IPD)?

UNDECIDED

Drug and device information, study documents

Studies a U.S. FDA-regulated drug product

No

Studies a U.S. FDA-regulated device product

No

This information was retrieved directly from the website clinicaltrials.gov without any changes. If you have any requests to change, remove or update your study details, please contact register@clinicaltrials.gov. As soon as a change is implemented on clinicaltrials.gov, this will be updated automatically on our website as well.

Subscribe