- ICH GCP
- US-Register für klinische Studien
- Klinische Studie NCT07739121
Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations (BEACON)
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
Studienübersicht
Status
Bedingungen
Intervention / Behandlung
Detaillierte Beschreibung
BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.
Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.
Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.
Studientyp
Einschreibung (Geschätzt)
Kontakte und Standorte
Studienkontakt
- Name: Jérôme Lambert, MD PhD
- Telefonnummer: +33 0142499742
- E-Mail: jerome.lambert@u-paris.fr
Studieren Sie die Kontaktsicherung
- Name: Jean-Emmanuel Bibault, MD PhD
- Telefonnummer: +33 01 56 09 34 06
- E-Mail: jean-emmanuel.bibault@aphp.fr
Studienorte
-
-
-
Paris, Frankreich
- Rekrutierung
- Hôpital Europeén Georges Pompidou
-
Kontakt:
- Jean-Emmanuel Bibault, MD PhD
- Telefonnummer: +33 01 56 09 28 34
- E-Mail: jean-emmanuel.bibault@aphp.fr
-
-
Teilnahmekriterien
Zulassungskriterien
Studienberechtigtes Alter
- Erwachsene
- Älterer Erwachsener
Akzeptiert gesunde Freiwillige
Probenahmeverfahren
Studienpopulation
Beschreibung
Inclusion Criteria:
- Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
- Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
- A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
Exclusion Criteria:
- Case outside the five predefined localisations.
- Incomplete, internally inconsistent or ambiguous schema.
- Duplicate or near-duplicate of an existing case in the set.
- Question not resolvable by current guidelines.
Studienplan
Wie ist die Studie aufgebaut?
Designdetails
Kohorten und Interventionen
Gruppe / Kohorte |
Intervention / Behandlung |
|---|---|
|
Breast cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Lung cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Urological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Digestive cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Gynaecological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
Was misst die Studie?
Primäre Ergebnismessungen
Ergebnis Maßnahme |
Maßnahmenbeschreibung |
Zeitfenster |
|---|---|---|
|
Domain-level performance between LLM recommendations and the locked guidelines.
Zeitfenster: Assessed once at central scoring, after data collection (~October 2026)
|
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
|
Assessed once at central scoring, after data collection (~October 2026)
|
Sekundäre Ergebnismessungen
Ergebnis Maßnahme |
Maßnahmenbeschreibung |
Zeitfenster |
|---|---|---|
|
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
Zeitfenster: Up to October 2026
|
Up to October 2026
|
|
|
Domain-level recommendation concordance between LLM and tumour-boards
Zeitfenster: Up to October 2026
|
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
|
Up to October 2026
|
|
Inter-tumour board domain-level recommendation concordance
Zeitfenster: Up to October 2026
|
Agreement between the two independent tumour boards scored per recommendation domain
|
Up to October 2026
|
|
Equipoise rate
Zeitfenster: Up to October 2026
|
Proportion of case-domains where the two tumour boards give different categorical recommendations
|
Up to October 2026
|
|
Completeness
Zeitfenster: Up to October 2026
|
Proportion of required domains addressed (LLM and tumour boards)
|
Up to October 2026
|
|
Missingness
Zeitfenster: Up to October 2026
|
Proportion of critical omissions (LLM and tumour boards)
|
Up to October 2026
|
|
Intensity bias
Zeitfenster: Up to October 2026
|
Proportion of recommendation corresponding to over- or under-treatment
|
Up to October 2026
|
Mitarbeiter und Ermittler
Studienaufzeichnungsdaten
Haupttermine studieren
Studienbeginn (Tatsächlich)
Primärer Abschluss (Geschätzt)
Studienabschluss (Geschätzt)
Studienanmeldedaten
Zuerst eingereicht
Zuerst eingereicht, das die QC-Kriterien erfüllt hat
Zuerst gepostet (Tatsächlich)
Studienaufzeichnungsaktualisierungen
Letztes Update gepostet (Tatsächlich)
Letztes eingereichtes Update, das die QC-Kriterien erfüllt
Zuletzt verifiziert
Mehr Informationen
Begriffe im Zusammenhang mit dieser Studie
Schlüsselwörter
- Neubildungen
- Patientensicherheit
- Künstliche Intelligenz
- Konkordanz
- Entscheidungsunterstützungssysteme, Klinisch
- Patientenbetreuungsteam
- Reproduzierbarkeit
- Benchmarking
- Große Sprachmodelle
- Benchmark
- Clinical Decision support
- Multidisciplinary tumour board (RCP)
- Medical oncology
- Weighted kappa
- Synthetic data
Zusätzliche relevante MeSH-Bedingungen
- Urogenitale Erkrankungen
- Genitalerkrankungen
- Genitale Neubildungen, männlich
- Urogenitale Neoplasmen
- Neubildungen nach Standort
- Genitalerkrankungen, männlich
- Prostataerkrankungen
- Männliche Urogenitalerkrankungen
- Nierenerkrankungen
- Urologische Erkrankungen
- Weibliche Urogenitalerkrankungen
- Weibliche Urogenitalerkrankungen und Schwangerschaftskomplikationen
- Erkrankungen der Atemwege
- Erkrankungen des Verdauungssystems
- Lungenkrankheit
- Neubildungen der Atemwege
- Thoraxneoplasmen
- Hautkrankheiten
- Brusterkrankungen
- Erkrankungen der Harnblase
- Haut- und Bindegewebserkrankungen
- Neubildungen
- Prostataneoplasmen
- Lungentumoren
- Neoplasien der Brust
- Neoplasien der Harnblase
- Urologische Neubildungen
- Nierentumoren
- Neoplasmen des Verdauungssystems
Andere Studien-ID-Nummern
- APHP261032
Plan für individuelle Teilnehmerdaten (IPD)
Planen Sie, individuelle Teilnehmerdaten (IPD) zu teilen?
Arzneimittel- und Geräteinformationen, Studienunterlagen
Studiert ein von der US-amerikanischen FDA reguliertes Arzneimittelprodukt
Studiert ein von der US-amerikanischen FDA reguliertes Geräteprodukt
Diese Informationen wurden ohne Änderungen direkt von der Website clinicaltrials.gov abgerufen. Wenn Sie Ihre Studiendaten ändern, entfernen oder aktualisieren möchten, wenden Sie sich bitte an register@clinicaltrials.gov. Sobald eine Änderung auf clinicaltrials.gov implementiert wird, wird diese automatisch auch auf unserer Website aktualisiert .