AI-Augmented Diagnostic Assessment With ENLIGHT Versus Independent Pathologist Review (ENLIGHT)

July 28, 2026 updated by: Kun-Hsing Yu, Harvard Medical School (HMS and HSDM)

This study will evaluate whether artificial intelligence (AI) can enhance clinicians' accuracy, efficiency, and confidence in distinguishing lung adenocarcinoma (LUAD) from lung squamous cell carcinoma (LUSC) and kidney renal papillary cell carcinoma (KIRP) from kidney renal clear cell carcinoma (KIRC) using digitized pathology slides. These subtype classifications are routinely performed by pathologists but can be challenging and time-consuming, particularly in difficult cases.

During the study, participating clinicians will review lung and kidney pathology slides under three different conditions:

  • Unaided Review: Diagnosis without AI assistance.
  • AI as Double-Check: The clinician first makes an independent diagnosis, after which the AI-generated diagnosis (prediction only or prediction with explanation) is revealed for review.
  • AI as First-Look: The AI-generated diagnosis (prediction only or prediction with explanation) is presented before the clinician begins the review.

Clinicians will be randomly assigned to different review sequences to minimize potential order effects. This study design will enable us to assess the impact of AI assistance on diagnostic accuracy, interpretation time, and clinician confidence.

Study Overview

Detailed Description

This study aims to evaluate the effect of artificial intelligence (AI) assistance on clinicians' diagnostic performance in distinguishing lung adenocarcinoma (LUAD) from lung squamous cell carcinoma (LUSC) and kidney renal papillary cell carcinoma (KIRP) from kidney renal clear cell carcinoma (KIRC) using digitized hematoxylin and eosin (H&E)-stained whole-slide images (WSIs). ENLIGHT (Explainable Neoplasm Learning In Grounded Histology Terms) will serve as the AI system under evaluation. This is a single-session, within-reader, between-case study in which each reader evaluates distinct sets of cases under all study conditions.

The study includes three diagnostic blocks: Block X, in which WSIs are reviewed without AI assistance; Block Y1, in which clinicians make an initial diagnosis before viewing the AI output as a double-check; and Block Y2, in which the AI output is displayed before clinicians begin their review as a first-look aid. Within each AI-assisted block, the prediction-only and prediction-with-explanation sub-blocks are presented in randomized order.

Each participating pathologist will review up to 400 de-identified WSIs (up to 200 lung cancer and up to 200 kidney cancer cases). Readers will be randomly assigned to one of four study arms that differ only in the order in which Blocks X, Y1, and Y2 are completed. For each reader, distinct WSIs will be randomly assigned to the diagnostic conditions so that no WSI is reviewed more than once by the same reader.

  • Arm 1 (X -> Y1 -> Y2): Clinicians first complete Block X (Unaided Review), followed by Block Y1 (AI as Double-Check) and then Block Y2 (AI as First-Look).
  • Arm 2 (X -> Y2 -> Y1): Clinicians first complete Block X (Unaided Review), followed by Block Y2 (AI as First-Look) and then Block Y1 (AI as Double-Check).
  • Arm 3 (Y1 -> Y2 -> X): Clinicians first complete Block Y1 (AI as Double-Check), followed by Block Y2 (AI as First-Look), and then Block X (Unaided Review).
  • Arm 4 (Y2 -> Y1 -> X): Clinicians first complete Block Y2 (AI as First-Look), followed by Block Y1 (AI as Double-Check), and then Block X (Unaided Review).

For each case, diagnostic accuracy, time to diagnosis, and diagnostic confidence will be recorded. No reader will review the same WSI under more than one condition, thereby eliminating within-reader recall bias. In parallel, the ENLIGHT model will independently generate diagnostic predictions for all WSIs to enable direct benchmarking of AI performance against pathologists and to evaluate the impact of different AI-assisted workflows on diagnostic performance.

Study Type

Interventional

Enrollment (Estimated)

25

Phase

  • Not Applicable

Contacts and Locations

This section provides the contact details for those conducting the study, and information on where this study is being conducted.

Study Locations

    • Massachusetts
      • Boston, Massachusetts, United States, 02115
        • Harvard Medical School,

Participation Criteria

Researchers look for people who fit a certain description, called eligibility criteria. Some examples of these criteria are a person's general health condition or prior treatments.

Eligibility Criteria

Ages Eligible for Study

  • Child
  • Adult
  • Older Adult

Accepts Healthy Volunteers

No

Description

Inclusion Criteria for Pathology Slides (i.e., Cases):

  • Hematoxylin and eosin (H&E)-stained pathology slides
  • Final diagnosis confirmed through molecular testing in conjunction with expert pathology evaluation

Exclusion Criteria for Pathology Slides (i.e., Cases):

  • Poor-quality or unreadable slides
  • Cases used in AI training

Inclusion Criteria for Readers (i.e., Participants):

  • Board-certified or board-eligible pathologists
  • Willingness to complete both unaided and AI-assisted review sessions

Study Plan

This section provides details of the study plan, including how the study is designed and what the study is measuring.

How is the study designed?

Design Details

  • Primary Purpose: Diagnostic
  • Allocation: Randomized
  • Interventional Model: Crossover Assignment
  • Masking: Quadruple

Arms and Interventions

Participant Group / Arm
Intervention / Treatment
Active Comparator: Unaided Review First, Then AI as Double-Check, Then AI as First-Look.
Readers first complete Block X (Unaided) on their assigned subset SX. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. For each reader, each of the five subsets (SX, SY1a, SY1b, SY2a, and SY2b) comprises up to 80 slides: up to 40 slides from LUAD-LUSC and up to 40 slides from KIRP-KIRC.
Readers first complete Block X (Unaided) on their assigned subset SX. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.
Active Comparator: Unaided Review First, Then AI as First-Look, Then AI as Double-Check.
Readers first complete Block X (Unaided) on their assigned subset SX. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. For each reader, each of the five subsets (SX, SY1a, SY1b, SY2a, and SY2b) comprises up to 80 slides: up to 40 slides from LUAD-LUSC and up to 40 slides from KIRP-KIRC.
Readers first complete Block X (Unaided) on their assigned subset SX. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.
Active Comparator: AI as Double-Check Review First, Then AI as First-Look, Then Unaided Review.
Readers first complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. Then readers complete Block X (Unaided) on their assigned subset SX. For each reader, each of the five subsets (SX, SY1a, SY1b, SY2a, and SY2b) comprises up to 80 slides: up to 40 slides from LUAD-LUSC and up to 40 slides from KIRP-KIRC.
Readers first complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. Then readers complete Block X (Unaided) on their assigned subset SX. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.
Active Comparator: AI as First-Look Review First, Then AI as Double-Check, Then Unaided Review.
Readers first complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. Then readers complete Block X (Unaided) on their assigned subset SX. For each reader, each of the five subsets (SX, SY1a, SY1b, SY2a, and SY2b) comprises up to 80 slides: up to 40 slides from LUAD-LUSC and up to 40 slides from KIRP-KIRC.
Readers first complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. Then readers complete Block X (Unaided) on their assigned subset SX. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.

What is the study measuring?

Primary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Diagnostic performance of cancers
Time Frame: Periprocedural (at the time of slide review)
Performance of clinicians (unaided and AI-assisted) for distinguishing LUAD- LUSC and distinguishing KIRP-KIRC, measured in accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1.
Periprocedural (at the time of slide review)

Secondary Outcome Measures

Outcome Measure
Measure Description
Time Frame
Time to diagnosis
Time Frame: Periprocedural (at the time of slide review)
Average time (seconds per case) required to finalize a diagnosis.
Periprocedural (at the time of slide review)
Inter-observer variability
Time Frame: Periprocedural (at the time of slide review)
Agreement among clinicians across conditions, measured using inter-rater reliability metrics (e.g., kappa statistics).
Periprocedural (at the time of slide review)
Net benefit after AI exposure
Time Frame: Periprocedural (at the time of slide review)
The overall change in diagnostic accuracy attributable to AI assistance.
Periprocedural (at the time of slide review)
Clinician confidence level
Time Frame: Periprocedural (at the time of slide review)
Self-reported diagnostic confidence recorded for each case. Scale: 5 - Absolutely Certain; 4 - Mostly Certain; 3 - Unsure; 2 - Very Doubtful; 1 - Random Guess; With 5 being the highest confidence score and 1 being the lowest.
Periprocedural (at the time of slide review)

Collaborators and Investigators

This is where you will find people and organizations involved with this study.

Study record dates

These dates track the progress of study record and summary results submissions to ClinicalTrials.gov. Study records and reported results are reviewed by the National Library of Medicine (NLM) to make sure they meet specific quality control standards before being posted on the public website.

Study Major Dates

Study Start (Estimated)

July 1, 2026

Primary Completion (Estimated)

August 1, 2026

Study Completion (Estimated)

August 1, 2026

Study Registration Dates

First Submitted

July 28, 2026

First Submitted That Met QC Criteria

July 28, 2026

First Posted (Actual)

August 3, 2026

Study Record Updates

Last Update Posted (Actual)

August 3, 2026

Last Update Submitted That Met QC Criteria

July 28, 2026

Last Verified

July 1, 2026

More Information

Terms related to this study

Plan for Individual participant data (IPD)

Plan to Share Individual Participant Data (IPD)?

NO

Drug and device information, study documents

Studies a U.S. FDA-regulated drug product

No

Studies a U.S. FDA-regulated device product

No

This information was retrieved directly from the website clinicaltrials.gov without any changes. If you have any requests to change, remove or update your study details, please contact register@clinicaltrials.gov. As soon as a change is implemented on clinicaltrials.gov, this will be updated automatically on our website as well.

Subscribe