- ICH GCP
- 미국 임상 시험 레지스트리
- 임상시험 NCT07739121
Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations (BEACON)
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
연구 개요
상태
상세 설명
BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.
Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.
Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.
연구 유형
등록 (추정된)
연락처 및 위치
연구 연락처
- 이름: Jérôme Lambert, MD PhD
- 전화번호: +33 0142499742
- 이메일: jerome.lambert@u-paris.fr
연구 연락처 백업
- 이름: Jean-Emmanuel Bibault, MD PhD
- 전화번호: +33 01 56 09 34 06
- 이메일: jean-emmanuel.bibault@aphp.fr
연구 장소
-
-
-
Paris, 프랑스
- 모병
- Hôpital Europeén Georges Pompidou
-
연락하다:
- Jean-Emmanuel Bibault, MD PhD
- 전화번호: +33 01 56 09 28 34
- 이메일: jean-emmanuel.bibault@aphp.fr
-
-
참여기준
자격 기준
공부할 수 있는 나이
- 성인
- 고령자
건강한 자원 봉사자를 받아들입니다
샘플링 방법
연구 인구
설명
Inclusion Criteria:
- Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
- Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
- A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
Exclusion Criteria:
- Case outside the five predefined localisations.
- Incomplete, internally inconsistent or ambiguous schema.
- Duplicate or near-duplicate of an existing case in the set.
- Question not resolvable by current guidelines.
공부 계획
연구는 어떻게 설계됩니까?
디자인 세부사항
코호트 및 개입
그룹/코호트 |
개입 / 치료 |
|---|---|
|
Breast cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Lung cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Urological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Digestive cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
|
Gynaecological cancers
|
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case.
Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6,
Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
|
연구는 무엇을 측정합니까?
주요 결과 측정
결과 측정 |
측정값 설명 |
기간 |
|---|---|---|
|
Domain-level performance between LLM recommendations and the locked guidelines.
기간: Assessed once at central scoring, after data collection (~October 2026)
|
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
|
Assessed once at central scoring, after data collection (~October 2026)
|
2차 결과 측정
결과 측정 |
측정값 설명 |
기간 |
|---|---|---|
|
Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)
기간: Up to October 2026
|
Up to October 2026
|
|
|
Domain-level recommendation concordance between LLM and tumour-boards
기간: Up to October 2026
|
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
|
Up to October 2026
|
|
Inter-tumour board domain-level recommendation concordance
기간: Up to October 2026
|
Agreement between the two independent tumour boards scored per recommendation domain
|
Up to October 2026
|
|
Equipoise rate
기간: Up to October 2026
|
Proportion of case-domains where the two tumour boards give different categorical recommendations
|
Up to October 2026
|
|
Completeness
기간: Up to October 2026
|
Proportion of required domains addressed (LLM and tumour boards)
|
Up to October 2026
|
|
Missingness
기간: Up to October 2026
|
Proportion of critical omissions (LLM and tumour boards)
|
Up to October 2026
|
|
Intensity bias
기간: Up to October 2026
|
Proportion of recommendation corresponding to over- or under-treatment
|
Up to October 2026
|
공동 작업자 및 조사자
연구 기록 날짜
연구 주요 날짜
연구 시작 (실제)
기본 완료 (추정된)
연구 완료 (추정된)
연구 등록 날짜
최초 제출
QC 기준을 충족하는 최초 제출
처음 게시됨 (실제)
연구 기록 업데이트
마지막 업데이트 게시됨 (실제)
QC 기준을 충족하는 마지막 업데이트 제출
마지막으로 확인됨
추가 정보
이 연구와 관련된 용어
키워드
추가 관련 MeSH 약관
기타 연구 ID 번호
- APHP261032
개별 참가자 데이터(IPD) 계획
개별 참가자 데이터(IPD)를 공유할 계획입니까?
약물 및 장치 정보, 연구 문서
미국 FDA 규제 의약품 연구
미국 FDA 규제 기기 제품 연구
이 정보는 변경 없이 clinicaltrials.gov 웹사이트에서 직접 가져온 것입니다. 귀하의 연구 세부 정보를 변경, 제거 또는 업데이트하도록 요청하는 경우 register@clinicaltrials.gov. 문의하십시오. 변경 사항이 clinicaltrials.gov에 구현되는 즉시 저희 웹사이트에도 자동으로 업데이트됩니다. .