このページは自動翻訳されたものであり、翻訳の正確性は保証されていません。を参照してください。 英語版 ソーステキスト用。

Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations (BEACON)

2026年7月28日 更新者:Assistance Publique - Hôpitaux de Paris

Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning

BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.

調査の概要

詳細な説明

BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.

Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.

Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.

研究の種類

観察的

入学 (推定)

100

連絡先と場所

このセクションには、調査を実施する担当者の連絡先の詳細と、この調査が実施されている場所に関する情報が記載されています。

研究連絡先

研究連絡先のバックアップ

研究場所

参加基準

研究者は、適格基準と呼ばれる特定の説明に適合する人を探します。これらの基準のいくつかの例は、人の一般的な健康状態または以前の治療です。

適格基準

就学可能な年齢

  • 大人
  • 高齢者

健康ボランティアの受け入れ

いいえ

サンプリング方法

非確率サンプル

調査対象母集団

100 synthetic oncology treatment-planning cases (20 per localisation) across five localisations: breast, lung, urological (prostate, bladder / upper-tract urothelial, kidney), digestive and gynaecological. Each case is a structured JSON input specifying UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question. No human participants, no patient data and no identifiable individuals. Recommendations are produced by two independent tumour boards per localisation and by five frontier LLMs (queried May 2026).

説明

Inclusion Criteria:

  • Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
  • Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
  • A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.

Exclusion Criteria:

  • Case outside the five predefined localisations.
  • Incomplete, internally inconsistent or ambiguous schema.
  • Duplicate or near-duplicate of an existing case in the set.
  • Question not resolvable by current guidelines.

研究計画

このセクションでは、研究がどのように設計され、研究が何を測定しているかなど、研究計画の詳細を提供します。

研究はどのように設計されていますか?

デザインの詳細

コホートと介入

グループ/コホート
介入・治療
Breast cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Lung cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Urological cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Digestive cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Gynaecological cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.

この研究は何を測定していますか?

主要な結果の測定

結果測定
メジャーの説明
時間枠
Domain-level performance between LLM recommendations and the locked guidelines.
時間枠:Assessed once at central scoring, after data collection (~October 2026)
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
Assessed once at central scoring, after data collection (~October 2026)

二次結果の測定

結果測定
メジャーの説明
時間枠
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
時間枠:Up to October 2026
Up to October 2026
Domain-level recommendation concordance between LLM and tumour-boards
時間枠:Up to October 2026
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
Up to October 2026
Inter-tumour board domain-level recommendation concordance
時間枠:Up to October 2026
Agreement between the two independent tumour boards scored per recommendation domain
Up to October 2026
Equipoise rate
時間枠:Up to October 2026
Proportion of case-domains where the two tumour boards give different categorical recommendations
Up to October 2026
Completeness
時間枠:Up to October 2026
Proportion of required domains addressed (LLM and tumour boards)
Up to October 2026
Missingness
時間枠:Up to October 2026
Proportion of critical omissions (LLM and tumour boards)
Up to October 2026
Intensity bias
時間枠:Up to October 2026
Proportion of recommendation corresponding to over- or under-treatment
Up to October 2026

協力者と研究者

ここでは、この調査に関係する人々や組織を見つけることができます。

研究記録日

これらの日付は、ClinicalTrials.gov への研究記録と要約結果の提出の進捗状況を追跡します。研究記録と報告された結果は、国立医学図書館 (NLM) によって審査され、公開 Web サイトに掲載される前に、特定の品質管理基準を満たしていることが確認されます。

主要日程の研究

研究開始 (実際)

2026年5月1日

一次修了 (推定)

2026年10月1日

研究の完了 (推定)

2026年10月1日

試験登録日

最初に提出

2026年7月28日

QC基準を満たした最初の提出物

2026年7月28日

最初の投稿 (実際)

2026年7月31日

学習記録の更新

投稿された最後の更新 (実際)

2026年7月31日

QC基準を満たした最後の更新が送信されました

2026年7月28日

最終確認日

2026年7月1日

詳しくは

この情報は、Web サイト clinicaltrials.gov から変更なしで直接取得したものです。研究の詳細を変更、削除、または更新するリクエストがある場合は、register@clinicaltrials.gov。 までご連絡ください。 clinicaltrials.gov に変更が加えられるとすぐに、ウェブサイトでも自動的に更新されます。

購読する