Diagnostic Accuracy of GPT-4o and Claude for HEART Score Calculation in Chest Pain

Recruiting now

Conditions studied: Emergency Medicine, Artificial Intelligence (AI), Artificial Intelligence (AI) in Diagnosis, Chest Pain Rule Out Myocardial Infarction

In brief

This prospective observational diagnostic accuracy study evaluates whether large language models (LLMs) - GPT-4o (OpenAI, gpt-4o-2024-11-20) and Claude (Anthropic, claude-sonnet-4-6) - can accurately calculate HEART scores from unstructured Turkish clinical notes and predict 30-day major adverse cardiac events (MACE) in emergency department patients presenting with non-traumatic chest pain. The study will enroll 600 consecutive adult patients. For each patient, the same anonymized data (free-text anamnesis, ECG report text, troponin value, and age) will be independently processed by both LLMs via separate API calls with deterministic settings (temperature=0, JSON format). A three-expert consensus HEART score - derived through blinded independent scoring by three emergency medicine physicians with majority-vote adjudication - serves as the reference standard for agreement analysis. Actual 30-day MACE (all-cause death, AMI Type 1/2/4b, unplanned revascularization) determined via national health database and telephone follow-up serves as the outcome for diagnostic accuracy analysis. A secondary documentation-quality sub-study will quantify how spontaneously Turkish emergency anamnesis notes capture HEART score parameters.

Key facts

Study ID
NCT07626060
Run by
Marmara University Pendik Training and Research Hospital
People needed
690
Starts
2026-06-01
Expected to finish
2027-06-01
Last updated by the study team
2026-06-23

Who can join

Age: 18 and older. Sex: any. Healthy volunteers: not accepted.

You may qualify if…

You may not qualify if…

Where it is running

Full record on ClinicalTrials.gov

Trial information comes from ClinicalTrials.gov and is refreshed daily. TrialsForMe does not provide medical care and does not run the studies it lists.