A Preregistered Multi-Cohort Evaluation of the FATHOM AI System for Molecular Testing Prioritization to Support Clinical Trial Enrollment (FATHOM)
Many clinical trials evaluating cancer treatments require patients to undergo testing for specific molecular markers as part of eligibility screening, typically using immunohistochemistry or sequencing. Because relatively few patients may carry a required marker, trial investigators often test large numbers of patients to identify the few who may ultimately qualify for enrollment.
Pathology laboratories routinely produce hematoxylin-and-eosin (H&E) slides during cancer diagnosis. Pathology foundation models-large neural networks pretrained on millions of histology images-have shown promise in predicting molecular characteristics from these slides. Researchers can use these models to build classifiers that predict specific molecular markers and prioritize patients for confirmatory testing.
This study evaluates FATHOM (Facilitating Accrual through Tumor Histology and Omics Matching), an autonomous research system powered by large multimodal models. Its agents read registered clinical trial records, identify molecular markers used as enrollment criteria, build prediction models using pathology foundation models, select the individual models or model combinations that best meet prespecified criteria, set their decision thresholds, and determine whether to deploy them. Together, a prediction model, its decision threshold, and the decision to deploy it constitute an AI prediction policy.
Before FATHOM runs, the investigators preregister the clinical trial records that its agents may read, the cutoff date that defines which trial information they may use, the rules governing the agents, and the analysis plan. The system timestamps and locks each policy immediately after an agent produces it. The investigators then apply the policies to archived patient slides and compare their predictions with existing molecular marker results.
The primary outcome is the proportion of prespecified evaluation scenarios in which an agent-generated policy, compared with universal molecular testing, either enriches the population selected for confirmatory testing with marker-positive patients or safely spares patients from confirmatory testing while meeting prespecified performance criteria.
This study analyzes existing pathology images and clinical trial records only. It does not enroll or contact patients, influence patient care, or affect participation in any clinical trial.
研究概览
详细说明
WHAT IS REGISTERED. This study evaluates screening policies generated by autonomous research agents using archived pathology slides and existing molecular marker profiles. The study does not prospectively enroll or contact patients. The investigators register the clinical trial corpus that the agents may read, the trial-record cutoff date, the information that the agents may access, the rules governing their work, and the methods used to score their policies.
THE AGENTS. Each agent may access the registered clinical trial corpus, published literature, out-of-fold performance estimates for its own classifiers, and any development data identified in its manifest. The agent receives no results from any sealed evaluation cohort. For visual recognition, each agent uses pathology AI models as feature extractors and trains classifiers on the extracted features. The agent determines which molecular markers to model, which model to use, how to set each operating threshold, and whether to deploy the resulting policy. The agent records each decision and its rationale.
THE MANIFEST. Each agent run produces a timestamped manifest listing every policy generated and each policy's final deployment decision. On the study start date, the investigators designate the autonomous-agent approach and its comparators for the primary evaluation.
TRIAL DEMAND CUTOFF. Trials first posted before January 1, 2026, define retrospective trial demand. Trials first posted on or after January 1, 2026, are used for the temporal generalization evaluation.
SEALED ANALYSIS RULE. Each policy result corresponds to an evaluation scenario, defined as one cohort paired with one molecular marker. For each scenario, the study logs and publishes the date on which investigators first compare any model output with the ground-truth marker result.
COHORTS. The study uses archived institutional and consortium cohorts containing routine diagnostic H&E slides linked to molecular profiles. No clinician uses model output from this study to make patient-care decisions.
研究类型
注册 (估计的)
联系人和位置
学习联系方式
- 姓名:Chi-Kang Pai
- 电话号码:6172334924
- 邮箱:chikang_pai@fas.harvard.edu
学习地点
-
-
Massachusetts
-
Boston、Massachusetts、美国、02115
- Harvard Medical School
-
接触:
- Chi-Kang Pai
- 电话号码:6172334924
- 邮箱:chikang_pai@fas.harvard.edu
-
首席研究员:
- Kun-Hsing Yu
-
-
参与标准
资格标准
适合学习的年龄
- 孩子
- 成人
- 年长者
接受健康志愿者
取样方法
研究人群
描述
Inclusion Criteria:
- Patients with a histologically confirmed cancer
- Availability of relevant molecular profiling results
- At least one diagnostic hematoxylin and eosin (H&E) whole-slide image
Exclusion Criteria:
- Poor-quality or unreadable slides, assessed independently of model output
- Patients whose slides were used to train a policy's classifier, for that policy's evaluation
学习计划
研究是如何设计的?
设计细节
队列和干预
团体/队列 |
|---|
|
Archived evaluation cohorts
Patient records and data from archived multi-institutional cohorts with routine H&E whole-slide images and molecular profiles.
No intervention is assigned, and no patient is contacted.
|
研究衡量的是什么?
主要结果指标
结果测量 |
措施说明 |
大体时间 |
|---|---|---|
|
Proportion of prespecified evaluation scenarios in which an AI-generated deployment policy demonstrates effective screening enrichment or rule-out performance
大体时间:Periprocedural (at the time of pathology slide evaluation)
|
An evaluation scenario consists of one molecular marker evaluated in one study cohort.
For each prespecified scenario, the study assesses whether the AI system generates a policy that either prioritizes patients more likely to carry the marker for confirmatory testing or identifies patients who may safely be spared testing, compared with testing everyone.
The outcome is the proportion of scenarios in which the policy meets these performance criteria.
|
Periprocedural (at the time of pathology slide evaluation)
|
次要结果测量
结果测量 |
措施说明 |
大体时间 |
|---|---|---|
|
Per-scenario performance of each AI-generated deployment policy
大体时间:Periprocedural (at the time of pathology slide evaluation)
|
For each evaluation scenario, the study reports sensitivity, negative predictive value, positive predictive value, the proportion of patients spared confirmatory testing, and the applicable enrichment or depletion ratio with its confidence interval.
|
Periprocedural (at the time of pathology slide evaluation)
|
|
Temporal generalizability for trials first posted on or after January 1, 2026
大体时间:Periprocedural (at the time of pathology slide evaluation)
|
Among trials first posted on or after January 1, 2026 that require a molecular biomarker for enrollment, the study evaluates: (1) the proportion of biomarkers and trials for which the agentic AI screening approach is useful; and (2) the estimated number of patients who would benefit from AI-guided screening compared with universal molecular testing.
Estimates are based on model performance and biomarker prevalence observed in the evaluation cohorts.
|
Periprocedural (at the time of pathology slide evaluation)
|
|
Temporal performance for trials first posted on or before December 31, 2025
大体时间:Periprocedural (at the time of pathology slide evaluation)
|
Among trials first posted on or before December 31, 2025, that require a molecular biomarker for enrollment, the study evaluates: (1) the proportion of biomarkers and trials for which the agentic AI screening approach is useful; and (2) the estimated number of patients who would benefit from AI-guided screening compared with universal molecular testing.
Estimates are based on model performance and biomarker prevalence observed in the evaluation cohorts.
|
Periprocedural (at the time of pathology slide evaluation)
|
|
Proportion of evaluation scenarios in which a non-default AI-generated deployment policy demonstrates effective screening enrichment or rule-out performance
大体时间:Periprocedural (at the time of pathology slide evaluation)
|
Among scenarios in which the deployed policy is not the default strategy of testing everyone, the study assesses whether the AI system generates a policy that either prioritizes patients more likely to carry the marker for confirmatory testing or identifies patients who may safely be spared testing, compared with testing everyone.
The outcome is the proportion of these scenarios in which the policy meets these performance criteria.
|
Periprocedural (at the time of pathology slide evaluation)
|
合作者和调查者
研究记录日期
研究主要日期
学习开始 (估计的)
初级完成 (估计的)
研究完成 (估计的)
研究注册日期
首次提交
首先提交符合 QC 标准的
首次发布 (实际的)
研究记录更新
最后更新发布 (实际的)
上次提交的符合 QC 标准的更新
最后验证
更多信息
与本研究相关的术语
其他相关的 MeSH 术语
- 泌尿生殖系统疾病
- 生殖器疾病
- 内分泌系统疾病
- 泌尿生殖系统肿瘤
- 按部位分类的肿瘤
- 男性泌尿生殖系统疾病
- 肾脏疾病
- 泌尿系统疾病
- 女性泌尿生殖系统疾病
- 女性泌尿生殖系统疾病和妊娠并发症
- 肠道疾病
- 呼吸道疾病
- 组织学类型的肿瘤
- 消化道肿瘤
- 消化系统肿瘤
- 消化系统疾病
- 肠胃疾病
- 胃病
- 肠道肿瘤
- 直肠疾病
- 子宫疾病
- 生殖器疾病,女性
- 肺部疾病
- 内分泌腺肿瘤
- 胰腺疾病
- 肿瘤、腺体和上皮
- 呼吸道肿瘤
- 胸部肿瘤
- 结肠疾病
- 卵巢疾病
- 附件疾病
- 生殖器肿瘤,女性
- 性腺疾病
- 皮肤病
- 乳腺疾病
- 泌尿系肿瘤
- 肿瘤,神经上皮
- 神经外胚层肿瘤
- 肿瘤、生殖细胞和胚胎
- 肿瘤,神经组织
- 子宫肿瘤
- 皮肤和结缔组织疾病
- 肿瘤
- 胃肿瘤
- 肺肿瘤
- 结直肠肿瘤
- 卵巢肿瘤
- 乳腺肿瘤
- 胰腺肿瘤
- 胶质瘤
- 头颈肿瘤
- 子宫内膜肿瘤
- 肾脏肿瘤
其他研究编号
- FATHOM
计划个人参与者数据 (IPD)
计划共享个人参与者数据 (IPD)?
药物和器械信息、研究文件
研究美国 FDA 监管的药品
研究美国 FDA 监管的设备产品
此信息直接从 clinicaltrials.gov 网站检索,没有任何更改。如果您有任何更改、删除或更新研究详细信息的请求,请联系 register@clinicaltrials.gov. clinicaltrials.gov 上实施更改,我们的网站上也会自动更新.