此页面是自动翻译的,不保证翻译的准确性。请参阅 英文版 对于源文本。

Large Language Model-Assisted cTNM Annotation From Chinese PSMA PET/CT Reports (PSMA-LLM-cTNM)

2026年7月20日 更新者:Qi Lin, MD、First Affiliated Hospital of Wenzhou Medical University

Large Language Model-Assisted Imaging cTNM Staging Annotation and Uncertainty Recognition for Prostate Cancer Based on Chinese PSMA PET/CT Reports

This observational study will develop and validate a large language model-assisted workflow for imaging cTNM staging annotation and uncertainty recognition in prostate cancer using Chinese PSMA PET/CT report texts generated during routine clinical care. The study will use de-identified report texts and necessary baseline clinical information only. No additional imaging examination, blood test, treatment, or follow-up visit will be assigned for this study.

The main objective is to evaluate whether a locally or institutionally controlled large language model can identify report-derived imaging cT, cN, and cM categories, extract supporting evidence from the original report, and recognize uncertainty expressions. Model performance will be assessed using an internal independent validation set, external validation reports from two collaborating hospitals, and a prospective validation set of 100 consecutive routine PSMA PET/CT reports. A human-AI comparison will also be performed using physicians from urology and imaging-related specialties with different seniority levels.

研究概览

详细说明

This is a multicenter observational diagnostic accuracy validation study based on Chinese PSMA PET/CT report texts from patients with prostate cancer or suspected prostate cancer. The study is not designed to evaluate a drug, device, surgical procedure, or imaging intervention. PSMA PET/CT examinations will be performed as part of routine clinical care, and the study will only analyze de-identified report texts and necessary baseline information after the reports have been finalized.

The study consists of retrospective and prospective components. Retrospectively, approximately 4,000 PSMA PET/CT reports from the First Affiliated Hospital of Wenzhou Medical University will be systematically annotated to construct a research database. An internal independent validation set of 300 reports, not used for model development or prompt optimization, will be used to evaluate the performance of the large language model. The reference standard for this 300-report validation set will be established by two experienced urologists through joint annotation, with adjudication by a nuclear medicine expert when needed. External validation will be performed using 110 de-identified reports from the First Affiliated Hospital of Ningbo University and 102 de-identified reports from Liuzhou People's Hospital. In addition, after ethics approval, 100 consecutive routine PSMA PET/CT reports from the First Affiliated Hospital of Wenzhou Medical University will be prospectively included to evaluate the accuracy and operational stability of the frozen model and prompt versions.

The large language model workflow will be deployed locally or in an institutionally controlled environment. The model will be instructed to generate structured JSON outputs, including cT_report, cN_report, cM_report, cT_uncertain, cN_uncertain, cM_uncertain, evidence_T, evidence_N, and evidence_M. The target task is report-derived imaging cTNM staging annotation, not pathological TNM staging or overall AJCC stage grouping. The model output will be used only for research evaluation and methodological analysis and will not be used for clinical diagnosis, treatment decision-making, or patient notification.

A human-AI comparison will be conducted on the 300-report internal validation set. Eight human evaluators from urology and imaging-related specialties, including trainees, residents, attending physicians, and associate chief physicians, will independently annotate the reports before and after learning the annotation manual. Annotation time will be recorded for each round. The performance of human evaluators and the large language model will be compared against the expert consensus reference standard.

The main outcome will be the accuracy of the large language model in identifying cT, cN, and cM categories from Chinese PSMA PET/CT reports. Secondary outcomes will include precision, recall, F1-score, macro-F1, micro-F1, complete cTNM triplet matching rate, uncertainty recognition performance, evidence extraction quality, human-AI comparison results, annotation time, external validation performance, prospective validation performance, and error type distribution. Error analysis will focus on local tumor extent, regional versus non-regional lymph node boundaries, bone and visceral metastasis recognition, equivocal wording, treatment-related context, benign or inflammatory alternatives, and lesions not attributable to prostate cancer.

研究类型

观察性的

注册 (估计的)

4600

联系人和位置

本节提供了进行研究的人员的详细联系信息,以及有关进行该研究的地点的信息。

学习联系方式

研究联系人备份

学习地点

    • Zhejiang
      • Wenzhou、Zhejiang、中国
        • 招聘中
        • The First Affiliated Hospital of Wenzhou Medical University
        • 接触:
        • 接触:

参与标准

研究人员寻找符合特定描述的人,称为资格标准。这些标准的一些例子是一个人的一般健康状况或先前的治疗。

资格标准

适合学习的年龄

  • 成人
  • 年长者

接受健康志愿者

不

取样方法

非概率样本

研究人群

The study population consists of male patients aged 18 years or older with clinically diagnosed, pathologically diagnosed, or clinically suspected prostate cancer who underwent PSMA PET/CT as part of routine clinical care. The study will include de-identified Chinese PSMA PET/CT report texts from the First Affiliated Hospital of Wenzhou Medical University, the First Affiliated Hospital of Ningbo University, and Liuzhou People's Hospital, as well as a prospective set of 100 consecutive routine PSMA PET/CT reports from the First Affiliated Hospital of Wenzhou Medical University.

描述

Inclusion Criteria:

  1. Male patients aged 18 years or older.
  2. Patients with clinically diagnosed, pathologically diagnosed, or clinically suspected prostate cancer.
  3. Patients who underwent PSMA PET/CT for initial staging, recurrence assessment, treatment response evaluation, metastatic assessment, or other clinical purposes during routine care.
  4. Complete or basically complete Chinese PSMA PET/CT report text is available, including imaging findings and/or diagnostic impression.
  5. The report text contains information that can be used to evaluate at least one target field, such as local prostate lesion, regional lymph nodes, non-regional lymph nodes, bone metastasis, visceral metastasis, or uncertainty expressions.
  6. The research data can be de-identified and replaced by a study identification number before analysis.

Exclusion Criteria:

  1. PSMA PET/CT reports unrelated to prostate cancer, or reports clearly irrelevant to the research task.
  2. Reports with severely missing, unreadable, or unavailable main text, imaging findings, or diagnostic impression.
  3. Reports that cannot be adequately de-identified or contain residual direct personal identifiers that cannot be safely removed.
  4. Duplicate records, repeated exports of the same examination, or records for which the unique report version cannot be confirmed.
  5. Reports judged by the research team to be of insufficient quality for manual annotation, model evaluation, or statistical analysis.

学习计划

本节提供研究计划的详细信息,包括研究的设计方式和研究的衡量标准。

研究是如何设计的?

设计细节

队列和干预

团体/队列
干预/治疗
PSMA PET/CT Report Text Validation Cohort
Patients with prostate cancer or suspected prostate cancer who underwent PSMA PET/CT as part of routine clinical care. De-identified Chinese PSMA PET/CT report texts and necessary baseline information will be used for manual annotation, large language model-assisted imaging cTNM staging annotation, uncertainty recognition, internal validation, external validation, prospective validation, and human-AI comparison. No additional examination, treatment, or follow-up will be assigned for this study.
A locally or institutionally controlled large language model workflow will analyze de-identified Chinese PSMA PET/CT report texts and generate structured outputs for report-derived imaging cTNM staging annotation, uncertainty recognition, and supporting evidence extraction. This workflow is used only for research evaluation and methodological analysis. It will not assign any examination, treatment, medication, procedure, or follow-up to participants, and it will not guide clinical diagnosis or treatment decisions.

研究衡量的是什么?

主要结果指标

结果测量
措施说明
大体时间
Accuracy of LLM-Assisted Imaging cTNM Staging Annotation
大体时间:After freezing the model and prompt versions, through completion of internal, external, and prospective validation, up to 18 months
The primary outcome is the accuracy of the large language model in identifying report-derived imaging cT, cN, and cM categories from de-identified Chinese PSMA PET/CT report texts. The LLM-generated cT_report, cN_report, and cM_report will be compared with the expert consensus reference standard. Accuracy, precision, recall, F1-score, macro-F1, micro-F1, complete cTNM triplet matching rate, and confusion matrices will be calculated in the internal 300-report validation set, external validation sets, and prospective 100-report validation set.
After freezing the model and prompt versions, through completion of internal, external, and prospective validation, up to 18 months

次要结果测量

结果测量
措施说明
大体时间
Component-Level Accuracy of LLM-Based Uncertainty Recognition
大体时间:After freezing the model and prompt versions, through completion of all validation analyses, up to 18 months.

This outcome measures the component-level accuracy of the large language model in recognizing uncertainty labels for report-derived imaging cTNM staging. The LLM-generated cT_uncertain, cN_uncertain, and cM_uncertain labels will be compared with the expert consensus reference standard. Accuracy will be calculated as the number of correctly classified uncertainty labels divided by the total number of component-level uncertainty labels across all reports. The three uncertainty components will be aggregated into one percentage value.

中文对应

After freezing the model and prompt versions, through completion of all validation analyses, up to 18 months.
Complete cTNM Triplet Matching Rate for Human Evaluators and the LLM
大体时间:During pre-training and post-training human annotation rounds and LLM batch inference, up to 18 months.
This outcome measures the proportion of reports for which the complete report-derived imaging cTNM triplet assigned by human evaluators and by the large language model exactly matches the expert consensus reference standard. A report will be counted as correct only when all three components, cT_report, cN_report, and cM_report, are correct. The result will be reported as the percentage of reports with complete cTNM triplet agreement. Results will be summarized separately for pre-training human annotation, post-training human annotation, and LLM batch inference.
During pre-training and post-training human annotation rounds and LLM batch inference, up to 18 months.
Annotation Time per Report for Human Evaluators and the LLM
大体时间:During pre-training and post-training human annotation rounds and LLM batch inference, up to 18 months.
This outcome measures the mean time required to complete report-derived imaging cTNM annotation per report. For human evaluators, annotation time will be recorded during the annotation rounds and divided by the number of annotated reports. For the LLM workflow, batch inference time will be divided by the number of processed reports. Results will be reported separately for pre-training human annotation, post-training human annotation, and LLM batch inference.
During pre-training and post-training human annotation rounds and LLM batch inference, up to 18 months.

合作者和调查者

在这里您可以找到参与这项研究的人员和组织。

研究记录日期

这些日期跟踪向 ClinicalTrials.gov 提交研究记录和摘要结果的进度。研究记录和报告的结果由国家医学图书馆 (NLM) 审查,以确保它们在发布到公共网站之前符合特定的质量控制标准。

研究主要日期

学习开始 (实际的)

2026年6月17日

初级完成 (估计的)

2027年6月30日

研究完成 (估计的)

2027年12月31日

研究注册日期

首次提交

2026年6月30日

首先提交符合 QC 标准的

2026年7月15日

首次发布 (实际的)

2026年7月16日

研究记录更新

最后更新发布 (实际的)

2026年7月21日

上次提交的符合 QC 标准的更新

2026年7月20日

最后验证

2026年7月1日

更多信息

与本研究相关的术语

计划个人参与者数据 (IPD)

计划共享个人参与者数据 (IPD)?

不

IPD 计划说明

Individual participant-level data will not be shared. The study data consist of de-identified Chinese PSMA PET/CT report texts and necessary baseline clinical information generated during routine clinical care. Although direct identifiers will be removed, the free-text report data may still carry a potential risk of re-identification. Therefore, individual-level raw data will not be made publicly available. Study findings will be reported in aggregate form. De-identified summary data or analysis methods may be made available upon reasonable request and with approval from the ethics committee and the institutional data governance authority, when applicable.

药物和器械信息、研究文件

研究美国 FDA 监管的药品

不

研究美国 FDA 监管的设备产品

不

此信息直接从 clinicaltrials.gov 网站检索,没有任何更改。如果您有任何更改、删除或更新研究详细信息的请求,请联系 register@clinicaltrials.gov. clinicaltrials.gov 上实施更改,我们的网站上也会自动更新.

订阅