このページは自動翻訳されたものであり、翻訳の正確性は保証されていません。を参照してください。 英語版 ソーステキスト用。

Agreement Between Large Language Model-Generated Treatment Recommendations With Guideline-Based and Tumor Board Decisions in Gastrointestinal Cancer (KITuKo)

2026年5月14日 更新者:Rene Mantke、Medizinische Hochschule Brandenburg Theodor Fontane

Concordance of Large Language Model-Generated Treatment Recommendations With Multidisciplinary Tumor Board and Guideline-Based Decisions in Gastrointestinal Cancer: A Retrospective Cohort Study

The goal of this observational study is to learn whether a computer program can suggest cancer treatments that match expert recommendations for people with gastrointestinal cancer (cancer of the pancreas, stomach, or colon and rectum).

The main questions it aims to answer are:

  • Do the treatment suggestions from the computer program match current medical guidelines?
  • Do these suggestions match decisions made by a multidisciplinary tumor board (a team of cancer specialists)?

Researchers will review existing medical records from people who have already been treated for these cancers. They will enter key clinical information into a computer program that uses artificial intelligence (AI). The program will generate treatment suggestions for each case.

Researchers will then compare these suggestions with:

  • guideline-based treatment recommendations
  • decisions made by the tumor board

This study will help researchers understand whether AI tools could support doctors in making cancer treatment decisions in the future.

調査の概要

詳細な説明

Gastrointestinal cancers require complex treatment planning that often involves surgery, systemic therapy, and multidisciplinary coordination. Clinical decision-making is typically guided by evidence-based recommendations and discussed in multidisciplinary tumor boards. However, the increasing complexity of treatment strategies and guideline frameworks can make consistent and reproducible decision-making challenging in routine clinical practice.

Recent advances in artificial intelligence have enabled the development of large language models (LLMs) that can process structured clinical information and generate text-based recommendations. These systems may offer a scalable approach to support clinical workflows, but their ability to produce reliable and clinically appropriate treatment suggestions in oncology remains uncertain.

This study evaluates the performance of an LLM-based system in the context of gastrointestinal oncology using retrospectively collected clinical case data. Structured case summaries derived from routine clinical documentation are used as standardized input. The model generates treatment recommendations under controlled conditions, allowing systematic comparison with established clinical reference standards.

The analysis focuses on the level of agreement between model-generated recommendations and established decision-making frameworks. In addition, the study explores how model performance varies across different clinical scenarios, including varying levels of disease complexity. Particular attention is given to situations in which recommendations differ, in order to better understand potential limitations of the model and identify patterns that may be clinically relevant.

Furthermore, the study examines the consistency of model outputs when the same clinical information is processed multiple times. This provides insight into the stability and reproducibility of the system, which are important considerations for potential real-world use.

The findings of this study are intended to inform the potential role of LLM-based tools as supportive systems in clinical decision-making. The study does not evaluate clinical outcomes or patient benefit, but instead focuses on agreement with established standards and expert-driven decisions as an initial step in assessing feasibility and safety.

研究の種類

観察的

入学 (実際)

30

連絡先と場所

このセクションには、調査を実施する担当者の連絡先の詳細と、この調査が実施されている場所に関する情報が記載されています。

研究場所

    • Brandenburg
      • Brandenburg an der Havel、Brandenburg、ドイツ、14770
        • University Hospital Brandenburg

参加基準

研究者は、適格基準と呼ばれる特定の説明に適合する人を探します。これらの基準のいくつかの例は、人の一般的な健康状態または以前の治療です。

適格基準

就学可能な年齢

  • 大人
  • 高齢者

健康ボランティアの受け入れ

いいえ

サンプリング方法

非確率サンプル

調査対象母集団

The study population consists of adult patients with gastrointestinal adenocarcinoma treated at a tertiary care academic center in the Federal State of Brandenburg, Germany. The population is derived from routine clinical practice and includes patients whose cases were evaluated in a multidisciplinary tumor board.

説明

Inclusion Criteria:

  • Histologically confirmed pancreatic, gastric, or colorectal adenocarcinoma
  • Treatment discussed in a multidisciplinary tumor board

Exclusion Criteria:

  • Non-adenocarcinoma histology

研究計画

このセクションでは、研究がどのように設計され、研究が何を測定しているかなど、研究計画の詳細を提供します。

研究はどのように設計されていますか?

デザインの詳細

コホートと介入

グループ/コホート
介入・治療
結腸直腸がん
結腸直腸がん患者
Detailed treatment recommendation according to the official guideline of the Association of the Scientific Medical Societies in Germany (AWMF; Arbeitsgemeinschaft der Wissenschaftlichen Medizinischen Fachgesellschaften),
Structured clinical case summaries were analyzed by a GPT-4-class large language model to generate treatment recommendations.
Detailed treatment recommendation according to the case-specific postoperative tumor board review.
Pancreatic cancer
Patients with pancreatic cancer
Detailed treatment recommendation according to the official guideline of the Association of the Scientific Medical Societies in Germany (AWMF; Arbeitsgemeinschaft der Wissenschaftlichen Medizinischen Fachgesellschaften),
Structured clinical case summaries were analyzed by a GPT-4-class large language model to generate treatment recommendations.
Detailed treatment recommendation according to the case-specific postoperative tumor board review.
Gastric cancer
Patients with gastric cancer
Detailed treatment recommendation according to the official guideline of the Association of the Scientific Medical Societies in Germany (AWMF; Arbeitsgemeinschaft der Wissenschaftlichen Medizinischen Fachgesellschaften),
Structured clinical case summaries were analyzed by a GPT-4-class large language model to generate treatment recommendations.
Detailed treatment recommendation according to the case-specific postoperative tumor board review.

この研究は何を測定していますか?

主要な結果の測定

結果測定
メジャーの説明
時間枠
Concordance with guideline-based management
時間枠:At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery
Agreement between LLM-generated recommendations and AWMF guideline-supported treatment strategies
At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery

二次結果の測定

結果測定
メジャーの説明
時間枠
Concordance with multidisciplinary tumor board decisions
時間枠:At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery
Agreement between LLM-generated recommendations and tumor board treatment strategies
At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery
Reproducibility of LLM recommendations across repeated runs
時間枠:At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery
Structured clinical case vignettes were entered into ChatGPT using a standardized prompt template. To assess within-model reproducibility, each clinical vignette was analyzed in 3 independent model sessions performed on different days using identical clinical input.
At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery
Characterization of discordant recommendations (e.g., overtreatment, undertreatment)
時間枠:At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery

Overtreatment was defined as an LLM-generated recommendation exceeding the intensity of the reference recommendation.

Undertreatment was defined as omission of a recommended treatment or recommendation of a less intensive strategy.

At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgery

協力者と研究者

ここでは、この調査に関係する人々や組織を見つけることができます。

研究記録日

これらの日付は、ClinicalTrials.gov への研究記録と要約結果の提出の進捗状況を追跡します。研究記録と報告された結果は、国立医学図書館 (NLM) によって審査され、公開 Web サイトに掲載される前に、特定の品質管理基準を満たしていることが確認されます。

主要日程の研究

研究開始 (実際)

2025年1月1日

一次修了 (実際)

2026年1月1日

研究の完了 (実際)

2026年2月25日

試験登録日

最初に提出

2026年5月4日

QC基準を満たした最初の提出物

2026年5月14日

最初の投稿 (実際)

2026年5月18日

学習記録の更新

投稿された最後の更新 (実際)

2026年5月18日

QC基準を満たした最後の更新が送信されました

2026年5月14日

最終確認日

2026年5月1日

詳しくは

本研究に関する用語

個々の参加者データ (IPD) の計画

個々の参加者データ (IPD) を共有する予定はありますか?

いいえ

IPD プランの説明

Individual participant data will not be shared. The dataset consists of retrospective, pseudonymized clinical data from a single institution, and sharing is restricted due to data protection regulations and institutional policies.

医薬品およびデバイス情報、研究文書

米国FDA規制医薬品の研究

いいえ

米国FDA規制機器製品の研究

いいえ

この情報は、Web サイト clinicaltrials.gov から変更なしで直接取得したものです。研究の詳細を変更、削除、または更新するリクエストがある場合は、register@clinicaltrials.gov。 までご連絡ください。 clinicaltrials.gov に変更が加えられるとすぐに、ウェブサイトでも自動的に更新されます。

購読する