Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models
2026년 7월 23일 업데이트: Mariam Ahmed Hossam, Cairo University
Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models in Endodontics Using Expert Consensus as the Reference Standard
This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment.
The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs.
Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.
연구 개요
상태
상태
아직 모집하지 않음
정황
정황
개입 / 치료
개입 / 치료
상세 설명
This prospective, blinded diagnostic accuracy study aims to compare the diagnostic accuracy and endodontic case difficulty assessment of three LLMs-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-with expert consensus as the reference standard.
Adult patients requiring primary endodontic treatment or retreatment will undergo routine clinical examination, pulp sensibility testing, and periapical radiographic assessment.
Standardized clinical information and radiographs will be provided to each LLM and to four experienced endodontists independently.
Expert consensus, defined as agreement among at least three of four experts, will serve as the reference standard.
The primary outcomes are diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus for pulpal and periapical diagnosis and endodontic case difficulty assessment according to the American Association of Endodontists (AAE) criteria.
The findings will provide evidence regarding the reliability of LLMs as clinical decision-support tools in endodontics and their potential role in improving diagnostic consistency and treatment planning.
연구 유형
연구 유형
관찰
등록 (추정된)
등록
342
연락처 및 위치
이 섹션에서는 연구를 수행하는 사람들의 연락처 정보와 이 연구가 수행되는 장소에 대한 정보를 제공합니다.
연구 연락처
연구 연락처
- 이름: Mariam Ahmed Hossam
- 전화번호: 00201110913251
- 이메일: mariam.ahmed.hosam@dentistry.cu.edu.eg
참여기준
연구원은 적격성 기준이라는 특정 설명에 맞는 사람을 찾습니다. 이러한 기준의 몇 가지 예는 개인의 일반적인 건강 상태 또는 이전 치료입니다.
자격 기준
자격 기준
공부할 수 있는 나이
- 어린이
- 성인
- 고령자
건강한 자원 봉사자를 받아들입니다
예
샘플링 방법
확률 샘플
연구 인구
Systemically healthy adult patients (≥16 years) requiring primary endodontic treatment or retreatment at the Endodontic Department, Faculty of Dentistry, Cairo University, who provide informed consent.
설명
Inclusion Criteria:
- Age above 16 years old.
- Systemically healthy patient (ASA I or II).
- Requiring endodontic treatment or retreatment
- Patient's acceptance to participate in the study
Exclusion Criteria:
- Medically compromised patients.
- Pregnant women.
- Traumatic dental injuries
- Low quality periapical radiograph
공부 계획
이 섹션에서는 연구 설계 방법과 연구가 측정하는 내용을 포함하여 연구 계획에 대한 세부 정보를 제공합니다.
연구는 어떻게 설계됩니까?
디자인 세부사항
연구는 무엇을 측정합니까?
주요 결과 측정
주요 결과 측정
결과 측정 |
측정값 설명 |
기간 |
|---|---|---|
|
Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosis
기간: At baseline
|
Diagnostic performance of each large language model compared with the expert consensus reference standard.
Accuracy, sensitivity, specificity, and agreement (Cohen's kappa) will be calculated according to the American Association of Endodontists (AAE) diagnostic criteria
|
At baseline
|
2차 결과 측정
2차 결과 측정
결과 측정 |
측정값 설명 |
기간 |
|---|---|---|
|
Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessment
기간: At baseline
|
Agreement between each large language model and the expert consensus in classifying endodontic case difficulty according to the AAE Endodontic Case Difficulty Assessment Guidelines.
Overall accuracy and weighted Cohen's kappa will be calculated.
|
At baseline
|
공동 작업자 및 조사자
여기에서 이 연구와 관련된 사람과 조직을 찾을 수 있습니다.
연구 기록 날짜
이 날짜는 ClinicalTrials.gov에 대한 연구 기록 및 요약 결과 제출의 진행 상황을 추적합니다. 연구 기록 및 보고된 결과는 공개 웹사이트에 게시되기 전에 특정 품질 관리 기준을 충족하는지 확인하기 위해 국립 의학 도서관(NLM)에서 검토합니다.
연구 주요 날짜
연구 시작 (추정된)
연구 시작
2026년 9월 1일
기본 완료 (추정된)
기본 완료
2026년 12월 1일
연구 완료 (추정된)
연구 완료
2027년 1월 1일
연구 등록 날짜
최초 제출
최초 제출
2026년 7월 10일
QC 기준을 충족하는 최초 제출
QC 기준을 충족하는 최초 제출
2026년 7월 23일
처음 게시됨 (실제)
처음 게시됨
2026년 7월 29일
연구 기록 업데이트
마지막 업데이트 게시됨 (실제)
마지막 업데이트 게시됨
2026년 7월 29일
QC 기준을 충족하는 마지막 업데이트 제출
QC 기준을 충족하는 마지막 업데이트 제출
2026년 7월 23일
마지막으로 확인됨
마지막으로 확인됨
2026년 7월 1일
추가 정보
이 연구와 관련된 용어
기타 연구 ID 번호
기타 연구 ID 번호
- New ENDO7.1.1
개별 참가자 데이터(IPD) 계획
개별 참가자 데이터(IPD)를 공유할 계획입니까?
미정
약물 및 장치 정보, 연구 문서
미국 FDA 규제 의약품 연구
아니
미국 FDA 규제 기기 제품 연구
아니
이 정보는 변경 없이 clinicaltrials.gov 웹사이트에서 직접 가져온 것입니다. 귀하의 연구 세부 정보를 변경, 제거 또는 업데이트하도록 요청하는 경우 register@clinicaltrials.gov. 문의하십시오. 변경 사항이 clinicaltrials.gov에 구현되는 즉시 저희 웹사이트에도 자동으로 업데이트됩니다. .