- ICH GCP
- Реестр клинических исследований США
- Клиническое испытание NCT07739121
Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations (BEACON)
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
Обзор исследования
Статус
Условия
- Новообразования почек
- Урологические новообразования
- Новообразования пищеварительной системы
- Новообразования молочной железы
- Новообразования легких
- Новообразования предстательной железы
- Новообразования мочевого пузыря
- Системы поддержки принятия решений, клинические
- Принятие решения
- Генитальные новообразования
- Искусственный интеллект
- Большие языковые модели
Вмешательство/лечение
Подробное описание
BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.
Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.
Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.
Тип исследования
Регистрация (Оцененный)
Контакты и местонахождение
Контакты исследования
- Имя: Jérôme Lambert, MD PhD
- Номер телефона: +33 0142499742
- Электронная почта: jerome.lambert@u-paris.fr
Учебное резервное копирование контактов
- Имя: Jean-Emmanuel Bibault, MD PhD
- Номер телефона: +33 01 56 09 34 06
- Электронная почта: jean-emmanuel.bibault@aphp.fr
Места учебы
-
-
-
Paris, Франция
- Рекрутинг
- Hôpital Europeén Georges Pompidou
-
Контакт:
- Jean-Emmanuel Bibault, MD PhD
- Номер телефона: +33 01 56 09 28 34
- Электронная почта: jean-emmanuel.bibault@aphp.fr
-
-
Критерии участия
Критерии приемлемости
Возраст, подходящий для обучения
- Взрослый
- Пожилой взрослый
Принимает здоровых добровольцев
Метод выборки
Исследуемая популяция
Описание
Inclusion Criteria:
- Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
- Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
- A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
Exclusion Criteria:
- Case outside the five predefined localisations.
- Incomplete, internally inconsistent or ambiguous schema.
- Duplicate or near-duplicate of an existing case in the set.
- Question not resolvable by current guidelines.
Учебный план
Как устроено исследование?
Детали дизайна
Когорты и вмешательства
Группа / когорта |
Вмешательство/лечение |
|---|---|
|
Breast cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Lung cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Urological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Digestive cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Gynaecological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
Что измеряет исследование?
Первичные показатели результатов
Мера результата |
Мера Описание |
Временное ограничение |
|---|---|---|
|
Domain-level performance between LLM recommendations and the locked guidelines.
Временное ограничение: Assessed once at central scoring, after data collection (~October 2026)
|
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
|
Assessed once at central scoring, after data collection (~October 2026)
|
Вторичные показатели результатов
Мера результата |
Мера Описание |
Временное ограничение |
|---|---|---|
|
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
Временное ограничение: Up to October 2026
|
Up to October 2026
|
|
|
Domain-level recommendation concordance between LLM and tumour-boards
Временное ограничение: Up to October 2026
|
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
|
Up to October 2026
|
|
Inter-tumour board domain-level recommendation concordance
Временное ограничение: Up to October 2026
|
Agreement between the two independent tumour boards scored per recommendation domain
|
Up to October 2026
|
|
Equipoise rate
Временное ограничение: Up to October 2026
|
Proportion of case-domains where the two tumour boards give different categorical recommendations
|
Up to October 2026
|
|
Completeness
Временное ограничение: Up to October 2026
|
Proportion of required domains addressed (LLM and tumour boards)
|
Up to October 2026
|
|
Missingness
Временное ограничение: Up to October 2026
|
Proportion of critical omissions (LLM and tumour boards)
|
Up to October 2026
|
|
Intensity bias
Временное ограничение: Up to October 2026
|
Proportion of recommendation corresponding to over- or under-treatment
|
Up to October 2026
|
Соавторы и исследователи
Даты записи исследования
Изучение основных дат
Начало исследования (Действительный)
Первичное завершение (Оцененный)
Завершение исследования (Оцененный)
Даты регистрации исследования
Первый отправленный
Впервые представлено, что соответствует критериям контроля качества
Первый опубликованный (Действительный)
Обновления учебных записей
Последнее опубликованное обновление (Действительный)
Последнее отправленное обновление, отвечающее критериям контроля качества
Последняя проверка
Дополнительная информация
Термины, связанные с этим исследованием
Ключевые слова
- Новообразования
- Безопасность пациентов
- Искусственный интеллект
- Согласие
- Системы поддержки принятия решений, клинические
- Команда по уходу за пациентами
- Воспроизводимость
- Бенчмаркинг
- Большие языковые модели
- Бенчмарк
- Clinical Decision support
- Multidisciplinary tumour board (RCP)
- Medical oncology
- Weighted kappa
- Synthetic data
Дополнительные соответствующие термины MeSH
- Урогенитальные заболевания
- Генитальные заболевания
- Генитальные новообразования, мужчины
- Урогенитальные новообразования
- Новообразования по локализации
- Заболевания половых органов, мужчины
- Заболевания предстательной железы
- Мужские мочеполовые заболевания
- Заболевания почек
- Урологические заболевания
- Женские урогенитальные заболевания
- Женские мочеполовые заболевания и осложнения беременности
- Заболевания дыхательных путей
- Заболевания пищеварительной системы
- Легочные заболевания
- Новообразования дыхательных путей
- Грудные новообразования
- Кожные заболевания
- Заболевания груди
- Заболевания мочевого пузыря
- Заболевания кожи и соединительной ткани
- Новообразования
- Новообразования предстательной железы
- Новообразования легких
- Новообразования молочной железы
- Новообразования мочевого пузыря
- Урологические новообразования
- Новообразования почек
- Новообразования пищеварительной системы
Другие идентификационные номера исследования
- APHP261032
Планирование данных отдельных участников (IPD)
Планируете делиться данными об отдельных участниках (IPD)?
Информация о лекарствах и устройствах, исследовательские документы
Изучает лекарственный продукт, регулируемый FDA США.
Изучает продукт устройства, регулируемый Управлением по санитарному надзору за качеством пищевых продуктов и медикаментов США.
Эта информация была получена непосредственно с веб-сайта clinicaltrials.gov без каких-либо изменений. Если у вас есть запросы на изменение, удаление или обновление сведений об исследовании, обращайтесь по адресу register@clinicaltrials.gov. Как только изменение будет реализовано на clinicaltrials.gov, оно будет автоматически обновлено и на нашем веб-сайте. .