Эта страница была переведена автоматически, точность перевода не гарантируется. Пожалуйста, обратитесь к английской версии для исходного текста.

Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations (BEACON)

28 июля 2026 г. обновлено: Assistance Publique - Hôpitaux de Paris

Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning

BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.

Обзор исследования

Подробное описание

BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.

Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.

Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.

Тип исследования

Наблюдательный

Регистрация (Оцененный)

100

Контакты и местонахождение

В этом разделе приведены контактные данные лиц, проводящих исследование, и информация о том, где проводится это исследование.

Контакты исследования

  • Имя: Jérôme Lambert, MD PhD
  • Номер телефона: +33 0142499742
  • Электронная почта: jerome.lambert@u-paris.fr

Учебное резервное копирование контактов

  • Имя: Jean-Emmanuel Bibault, MD PhD
  • Номер телефона: +33 01 56 09 34 06
  • Электронная почта: jean-emmanuel.bibault@aphp.fr

Места учебы

      • Paris, Франция
        • Рекрутинг
        • Hôpital Europeén Georges Pompidou
        • Контакт:

Критерии участия

Исследователи ищут людей, которые соответствуют определенному описанию, называемому критериям приемлемости. Некоторыми примерами этих критериев являются общее состояние здоровья человека или предшествующее лечение.

Критерии приемлемости

Возраст, подходящий для обучения

  • Взрослый
  • Пожилой взрослый

Принимает здоровых добровольцев

Нет

Метод выборки

Невероятностная выборка

Исследуемая популяция

100 synthetic oncology treatment-planning cases (20 per localisation) across five localisations: breast, lung, urological (prostate, bladder / upper-tract urothelial, kidney), digestive and gynaecological. Each case is a structured JSON input specifying UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question. No human participants, no patient data and no identifiable individuals. Recommendations are produced by two independent tumour boards per localisation and by five frontier LLMs (queried May 2026).

Описание

Inclusion Criteria:

  • Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
  • Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
  • A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.

Exclusion Criteria:

  • Case outside the five predefined localisations.
  • Incomplete, internally inconsistent or ambiguous schema.
  • Duplicate or near-duplicate of an existing case in the set.
  • Question not resolvable by current guidelines.

Учебный план

В этом разделе представлена ​​подробная информация о плане исследования, в том числе о том, как планируется исследование и что оно измеряет.

Как устроено исследование?

Детали дизайна

Когорты и вмешательства

Группа / когорта
Вмешательство/лечение
Breast cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Lung cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Urological cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Digestive cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Gynaecological cancers
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.

Что измеряет исследование?

Первичные показатели результатов

Мера результата
Мера Описание
Временное ограничение
Domain-level performance between LLM recommendations and the locked guidelines.
Временное ограничение: Assessed once at central scoring, after data collection (~October 2026)
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
Assessed once at central scoring, after data collection (~October 2026)

Вторичные показатели результатов

Мера результата
Мера Описание
Временное ограничение
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
Временное ограничение: Up to October 2026
Up to October 2026
Domain-level recommendation concordance between LLM and tumour-boards
Временное ограничение: Up to October 2026
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
Up to October 2026
Inter-tumour board domain-level recommendation concordance
Временное ограничение: Up to October 2026
Agreement between the two independent tumour boards scored per recommendation domain
Up to October 2026
Equipoise rate
Временное ограничение: Up to October 2026
Proportion of case-domains where the two tumour boards give different categorical recommendations
Up to October 2026
Completeness
Временное ограничение: Up to October 2026
Proportion of required domains addressed (LLM and tumour boards)
Up to October 2026
Missingness
Временное ограничение: Up to October 2026
Proportion of critical omissions (LLM and tumour boards)
Up to October 2026
Intensity bias
Временное ограничение: Up to October 2026
Proportion of recommendation corresponding to over- or under-treatment
Up to October 2026

Соавторы и исследователи

Здесь вы найдете людей и организации, участвующие в этом исследовании.

Даты записи исследования

Эти даты отслеживают ход отправки отчетов об исследованиях и сводных результатов на сайт ClinicalTrials.gov. Записи исследований и сообщаемые результаты проверяются Национальной медицинской библиотекой (NLM), чтобы убедиться, что они соответствуют определенным стандартам контроля качества, прежде чем публиковать их на общедоступном веб-сайте.

Изучение основных дат

Начало исследования (Действительный)

1 мая 2026 г.

Первичное завершение (Оцененный)

1 октября 2026 г.

Завершение исследования (Оцененный)

1 октября 2026 г.

Даты регистрации исследования

Первый отправленный

28 июля 2026 г.

Впервые представлено, что соответствует критериям контроля качества

28 июля 2026 г.

Первый опубликованный (Действительный)

31 июля 2026 г.

Обновления учебных записей

Последнее опубликованное обновление (Действительный)

31 июля 2026 г.

Последнее отправленное обновление, отвечающее критериям контроля качества

28 июля 2026 г.

Последняя проверка

1 июля 2026 г.

Дополнительная информация

Термины, связанные с этим исследованием

Дополнительные соответствующие термины MeSH

Другие идентификационные номера исследования

  • APHP261032

Планирование данных отдельных участников (IPD)

Планируете делиться данными об отдельных участниках (IPD)?

НЕ РЕШЕНО

Информация о лекарствах и устройствах, исследовательские документы

Изучает лекарственный продукт, регулируемый FDA США.

Нет

Изучает продукт устройства, регулируемый Управлением по санитарному надзору за качеством пищевых продуктов и медикаментов США.

Нет

Эта информация была получена непосредственно с веб-сайта clinicaltrials.gov без каких-либо изменений. Если у вас есть запросы на изменение, удаление или обновление сведений об исследовании, обращайтесь по адресу register@clinicaltrials.gov. Как только изменение будет реализовано на clinicaltrials.gov, оно будет автоматически обновлено и на нашем веб-сайте. .

Подписаться