Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations
Recruiting now
Conditions studied: Breast Neoplasms, Lung Neoplasms, Urologic Neoplasms, Prostatic Neoplasms, Urinary Bladder Neoplasms, Kidney Neoplasms, Digestive System Neoplasms, Genital Neoplasms, Artifical Intelligence, Large Language Models, Decision Making, Decision Support Systems, Clinical
In brief
BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.
Key facts
- Study ID
- NCT07739121
- Run by
- Assistance Publique - Hôpitaux de Paris
- People needed
- 100
- Starts
- 2026-05-01
- Expected to finish
- 2026-10-01
- Last updated by the study team
- 2026-07-31
Who can join
Age: 18 and older. Sex: any. Healthy volunteers: not accepted.
You may qualify if…
- Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
- Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
- A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.
You may not qualify if…
- Case outside the five predefined localisations.
- Incomplete, internally inconsistent or ambiguous schema.
- Duplicate or near-duplicate of an existing case in the set.
- Question not resolvable by current guidelines.
Where it is running
- Hopital Européen Georges Pompidou — Paris, France (enrolling)
Full record on ClinicalTrials.gov
Trial information comes from ClinicalTrials.gov and is refreshed daily. TrialsForMe does not provide medical care and does not run the studies it lists.