Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations (BEACON)
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
調査の概要
状態
詳細な説明
BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.
Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.
Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.
研究の種類
入学 (推定)
連絡先と場所
研究連絡先
- 名前:Jérôme Lambert, MD PhD
- 電話番号:+33 0142499742
- メール:jerome.lambert@u-paris.fr
研究連絡先のバックアップ
- 名前:Jean-Emmanuel Bibault, MD PhD
- 電話番号:+33 01 56 09 34 06
- メール:jean-emmanuel.bibault@aphp.fr
研究場所
-
-
-
Paris、フランス
- 募集
- Hôpital Europeén Georges Pompidou
-
コンタクト:
- Jean-Emmanuel Bibault, MD PhD
- 電話番号:+33 01 56 09 28 34
- メール:jean-emmanuel.bibault@aphp.fr
-
-
参加基準
適格基準
就学可能な年齢
- 大人
- 高齢者
健康ボランティアの受け入れ
サンプリング方法
調査対象母集団
説明
Inclusion Criteria:
- Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
- Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
- A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
Exclusion Criteria:
- Case outside the five predefined localisations.
- Incomplete, internally inconsistent or ambiguous schema.
- Duplicate or near-duplicate of an existing case in the set.
- Question not resolvable by current guidelines.
研究計画
研究はどのように設計されていますか?
デザインの詳細
コホートと介入
グループ/コホート |
介入・治療 |
|---|---|
|
Breast cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Lung cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Urological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Digestive cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Gynaecological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
この研究は何を測定していますか?
主要な結果の測定
結果測定 |
メジャーの説明 |
時間枠 |
|---|---|---|
|
Domain-level performance between LLM recommendations and the locked guidelines.
時間枠:Assessed once at central scoring, after data collection (~October 2026)
|
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
|
Assessed once at central scoring, after data collection (~October 2026)
|
二次結果の測定
結果測定 |
メジャーの説明 |
時間枠 |
|---|---|---|
|
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
時間枠:Up to October 2026
|
Up to October 2026
|
|
|
Domain-level recommendation concordance between LLM and tumour-boards
時間枠:Up to October 2026
|
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
|
Up to October 2026
|
|
Inter-tumour board domain-level recommendation concordance
時間枠:Up to October 2026
|
Agreement between the two independent tumour boards scored per recommendation domain
|
Up to October 2026
|
|
Equipoise rate
時間枠:Up to October 2026
|
Proportion of case-domains where the two tumour boards give different categorical recommendations
|
Up to October 2026
|
|
Completeness
時間枠:Up to October 2026
|
Proportion of required domains addressed (LLM and tumour boards)
|
Up to October 2026
|
|
Missingness
時間枠:Up to October 2026
|
Proportion of critical omissions (LLM and tumour boards)
|
Up to October 2026
|
|
Intensity bias
時間枠:Up to October 2026
|
Proportion of recommendation corresponding to over- or under-treatment
|
Up to October 2026
|
協力者と研究者
研究記録日
主要日程の研究
研究開始 (実際)
一次修了 (推定)
研究の完了 (推定)
試験登録日
最初に提出
QC基準を満たした最初の提出物
最初の投稿 (実際)
学習記録の更新
投稿された最後の更新 (実際)
QC基準を満たした最後の更新が送信されました
最終確認日
詳しくは
本研究に関する用語
キーワード
追加の関連 MeSH 用語
その他の研究ID番号
- APHP261032
個々の参加者データ (IPD) の計画
個々の参加者データ (IPD) を共有する予定はありますか?
医薬品およびデバイス情報、研究文書
米国FDA規制医薬品の研究
米国FDA規制機器製品の研究
この情報は、Web サイト clinicaltrials.gov から変更なしで直接取得したものです。研究の詳細を変更、削除、または更新するリクエストがある場合は、register@clinicaltrials.gov。 までご連絡ください。 clinicaltrials.gov に変更が加えられるとすぐに、ウェブサイトでも自動的に更新されます。