- ICH GCP
- Registro de ensaios clínicos dos EUA
- Ensaio Clínico NCT07739121
Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations (BEACON)
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
Visão geral do estudo
Status
Condições
Intervenção / Tratamento
Descrição detalhada
BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.
Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.
Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.
Tipo de estudo
Inscrição (Estimado)
Contactos e Locais
Contato de estudo
- Nome: Jérôme Lambert, MD PhD
- Número de telefone: +33 0142499742
- E-mail: jerome.lambert@u-paris.fr
Estude backup de contato
- Nome: Jean-Emmanuel Bibault, MD PhD
- Número de telefone: +33 01 56 09 34 06
- E-mail: jean-emmanuel.bibault@aphp.fr
Locais de estudo
-
-
-
Paris, França
- Recrutamento
- Hôpital Europeén Georges Pompidou
-
Contato:
- Jean-Emmanuel Bibault, MD PhD
- Número de telefone: +33 01 56 09 28 34
- E-mail: jean-emmanuel.bibault@aphp.fr
-
-
Critérios de participação
Critérios de elegibilidade
Idades elegíveis para estudo
- Adulto
- Adulto mais velho
Aceita Voluntários Saudáveis
Método de amostragem
População do estudo
Descrição
Inclusion Criteria:
- Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
- Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
- A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
Exclusion Criteria:
- Case outside the five predefined localisations.
- Incomplete, internally inconsistent or ambiguous schema.
- Duplicate or near-duplicate of an existing case in the set.
- Question not resolvable by current guidelines.
Plano de estudo
Como o estudo é projetado?
Detalhes do projeto
Coortes e Intervenções
Grupo / Coorte |
Intervenção / Tratamento |
|---|---|
|
Breast cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Lung cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Urological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Digestive cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Gynaecological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
O que o estudo está medindo?
Medidas de resultados primários
Medida de resultado |
Descrição da medida |
Prazo |
|---|---|---|
|
Domain-level performance between LLM recommendations and the locked guidelines.
Prazo: Assessed once at central scoring, after data collection (~October 2026)
|
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
|
Assessed once at central scoring, after data collection (~October 2026)
|
Medidas de resultados secundários
Medida de resultado |
Descrição da medida |
Prazo |
|---|---|---|
|
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
Prazo: Up to October 2026
|
Up to October 2026
|
|
|
Domain-level recommendation concordance between LLM and tumour-boards
Prazo: Up to October 2026
|
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
|
Up to October 2026
|
|
Inter-tumour board domain-level recommendation concordance
Prazo: Up to October 2026
|
Agreement between the two independent tumour boards scored per recommendation domain
|
Up to October 2026
|
|
Equipoise rate
Prazo: Up to October 2026
|
Proportion of case-domains where the two tumour boards give different categorical recommendations
|
Up to October 2026
|
|
Completeness
Prazo: Up to October 2026
|
Proportion of required domains addressed (LLM and tumour boards)
|
Up to October 2026
|
|
Missingness
Prazo: Up to October 2026
|
Proportion of critical omissions (LLM and tumour boards)
|
Up to October 2026
|
|
Intensity bias
Prazo: Up to October 2026
|
Proportion of recommendation corresponding to over- or under-treatment
|
Up to October 2026
|
Colaboradores e Investigadores
Patrocinador
Datas de registro do estudo
Datas Principais do Estudo
Início do estudo (Real)
Conclusão Primária (Estimado)
Conclusão do estudo (Estimado)
Datas de inscrição no estudo
Enviado pela primeira vez
Enviado pela primeira vez que atendeu aos critérios de CQ
Primeira postagem (Real)
Atualizações de registro de estudo
Última Atualização Postada (Real)
Última atualização enviada que atendeu aos critérios de controle de qualidade
Última verificação
Mais Informações
Termos relacionados a este estudo
Palavras-chave
- Neoplasias
- Segurança do paciente
- Inteligência artificial
- Concordância
- Sistemas de Apoio à Decisão, Clínica
- Equipe de atendimento ao paciente
- Reprodutibilidade
- Avaliação comparativa
- Grandes modelos de linguagem
- Referência
- Clinical Decision support
- Multidisciplinary tumour board (RCP)
- Medical oncology
- Weighted kappa
- Synthetic data
Termos MeSH relevantes adicionais
- Doenças urogenitais
- Doenças Genitais
- Neoplasias Genitais Masculinas
- Neoplasias urogenitais
- Neoplasias por local
- Doenças Genitais, Masculino
- Doenças prostáticas
- Doenças Urogenitais Masculinas
- Doenças renais
- Doenças Urológicas
- Doenças Urogenitais Femininas
- Doenças urogenitais femininas e complicações na gravidez
- Doenças Respiratórias
- Doenças do aparelho digestivo
- Doenças pulmonares
- Neoplasias do Trato Respiratório
- Neoplasias Torácicas
- Doenças de pele
- Doenças da mama
- Doenças da Bexiga Urinária
- Doenças da Pele e do Tecido Conjuntivo
- Neoplasias
- Neoplasias prostáticas
- Neoplasias Pulmonares
- Neoplasias da Mama
- Neoplasias da Bexiga Urinária
- Neoplasias Urológicas
- Neoplasias Renais
- Neoplasias do Aparelho Digestivo
Outros números de identificação do estudo
- APHP261032
Plano para dados de participantes individuais (IPD)
Planeja compartilhar dados de participantes individuais (IPD)?
Informações sobre medicamentos e dispositivos, documentos de estudo
Estuda um medicamento regulamentado pela FDA dos EUA
Estuda um produto de dispositivo regulamentado pela FDA dos EUA
Essas informações foram obtidas diretamente do site clinicaltrials.gov sem nenhuma alteração. Se você tiver alguma solicitação para alterar, remover ou atualizar os detalhes do seu estudo, entre em contato com register@clinicaltrials.gov. Assim que uma alteração for implementada em clinicaltrials.gov, ela também será atualizada automaticamente em nosso site .