Denna sida har översatts automatiskt och översättningens korrekthet kan inte garanteras. Vänligen se engelsk version för en källtext.

Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models

23 juli 2026 uppdaterad av: Mariam Ahmed Hossam, Cairo University

Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models in Endodontics Using Expert Consensus as the Reference Standard

This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.

Studieöversikt

Status

Har inte rekryterat ännu

Detaljerad beskrivning

This prospective, blinded diagnostic accuracy study aims to compare the diagnostic accuracy and endodontic case difficulty assessment of three LLMs-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-with expert consensus as the reference standard. Adult patients requiring primary endodontic treatment or retreatment will undergo routine clinical examination, pulp sensibility testing, and periapical radiographic assessment. Standardized clinical information and radiographs will be provided to each LLM and to four experienced endodontists independently. Expert consensus, defined as agreement among at least three of four experts, will serve as the reference standard. The primary outcomes are diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus for pulpal and periapical diagnosis and endodontic case difficulty assessment according to the American Association of Endodontists (AAE) criteria. The findings will provide evidence regarding the reliability of LLMs as clinical decision-support tools in endodontics and their potential role in improving diagnostic consistency and treatment planning.

Studietyp

Observationell

Inskrivning (Beräknad)

342

Kontakter och platser

Det här avsnittet innehåller kontaktuppgifter för dem som genomför studien och information om var denna studie genomförs.

Studiekontakt

Deltagandekriterier

Forskare letar efter personer som passar en viss beskrivning, så kallade behörighetskriterier. Några exempel på dessa kriterier är en persons allmänna hälsotillstånd eller tidigare behandlingar.

Urvalskriterier

Åldrar som är berättigade till studier

  • Barn
  • Vuxen
  • Äldre vuxen

Tar emot friska volontärer

Ja

Testmetod

Sannolikhetsprov

Studera befolkning

Systemically healthy adult patients (≥16 years) requiring primary endodontic treatment or retreatment at the Endodontic Department, Faculty of Dentistry, Cairo University, who provide informed consent.

Beskrivning

Inclusion Criteria:

  1. Age above 16 years old.
  2. Systemically healthy patient (ASA I or II).
  3. Requiring endodontic treatment or retreatment
  4. Patient's acceptance to participate in the study

Exclusion Criteria:

  1. Medically compromised patients.
  2. Pregnant women.
  3. Traumatic dental injuries
  4. Low quality periapical radiograph

Studieplan

Det här avsnittet ger detaljer om studieplanen, inklusive hur studien är utformad och vad studien mäter.

Hur är studien utformad?

Designdetaljer

Vad mäter studien?

Primära resultatmått

Resultatmått
Åtgärdsbeskrivning
Tidsram
Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosis
Tidsram: At baseline
Diagnostic performance of each large language model compared with the expert consensus reference standard. Accuracy, sensitivity, specificity, and agreement (Cohen's kappa) will be calculated according to the American Association of Endodontists (AAE) diagnostic criteria
At baseline

Sekundära resultatmått

Resultatmått
Åtgärdsbeskrivning
Tidsram
Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessment
Tidsram: At baseline
Agreement between each large language model and the expert consensus in classifying endodontic case difficulty according to the AAE Endodontic Case Difficulty Assessment Guidelines. Overall accuracy and weighted Cohen's kappa will be calculated.
At baseline

Samarbetspartners och utredare

Det är här du hittar personer och organisationer som är involverade i denna studie.

Studieavstämningsdatum

Dessa datum spårar framstegen för inlämningar av studieposter och sammanfattande resultat till ClinicalTrials.gov. Studieposter och rapporterade resultat granskas av National Library of Medicine (NLM) för att säkerställa att de uppfyller specifika kvalitetskontrollstandarder innan de publiceras på den offentliga webbplatsen.

Studera stora datum

Studiestart (Beräknad)

1 september 2026

Primärt slutförande (Beräknad)

1 december 2026

Avslutad studie (Beräknad)

1 januari 2027

Studieregistreringsdatum

Först inskickad

10 juli 2026

Först inskickad som uppfyllde QC-kriterierna

23 juli 2026

Första postat (Faktisk)

29 juli 2026

Uppdateringar av studier

Senaste uppdatering publicerad (Faktisk)

29 juli 2026

Senaste inskickade uppdateringen som uppfyllde QC-kriterierna

23 juli 2026

Senast verifierad

1 juli 2026

Mer information

Termer relaterade till denna studie

Andra studie-ID-nummer

  • New ENDO7.1.1

Plan för individuella deltagardata (IPD)

Planerar du att dela individuella deltagardata (IPD)?

OBESLUTSAM

Läkemedels- och apparatinformation, studiedokument

Studerar en amerikansk FDA-reglerad läkemedelsprodukt

Nej

Studerar en amerikansk FDA-reglerad produktprodukt

Nej

Denna information hämtades direkt från webbplatsen clinicaltrials.gov utan några ändringar. Om du har några önskemål om att ändra, ta bort eller uppdatera dina studieuppgifter, vänligen kontakta register@clinicaltrials.gov. Så snart en ändring har implementerats på clinicaltrials.gov, kommer denna att uppdateras automatiskt även på vår webbplats .

Prenumerera