Diagnostic Accuracy of GPT-4o and Claude 4.6 Sonnet in Turkish ED Anamnesis Notes
Recruiting now
Conditions studied: Emergency Medicine, Diagnostic Errors, Artificial Intelligence (AI) in Diagnosis
In brief
This retrospective diagnostic accuracy study evaluates the ability of two large language models (LLMs) - GPT-4o (gpt-4o-2024-11-20; OpenAI) and Claude 4.6 Sonnet (claude-sonnet-4-6; Anthropic) - to generate correct diagnoses from anonymized Turkish-language emergency department (ED) anamnesis notes, and compares their performance with the diagnosis entered by the treating emergency physician. A consensus gold standard is established by three independent board-certified emergency medicine specialists who blindly review each note and vote on the primary diagnosis using ICD-10 three-character codes; the majority vote (at least 2 of 3 specialists agreeing) constitutes the reference standard. Both LLMs are evaluated using a standardized zero-shot direct prompting strategy (temperature=0, stateless API sessions). The primary outcome is diagnostic accuracy (proportion of ICD-10 chapter-level matches) and Cohen's kappa for each LLM against the gold standard. Secondary outcomes include top-3 accuracy, treating physician accuracy, inter-model agreement, and subgroup analyses by ESI triage level and ICD-10 chapter. Inter-rater reliability among the three specialists is quantified using Fleiss' kappa. Analyses are performed in Jamovi. This study represents the first evaluation of LLM diagnostic accuracy using Turkish-language clinical notes and the first to benchmark LLM performance against an independent three-specialist majority-vote gold standard rather than against the treating physician's own diagnosis.
Key facts
- Study ID
- NCT07632859
- Run by
- Marmara University Pendik Training and Research Hospital
- People needed
- 600
- Starts
- 2026-06-01
- Expected to finish
- 2026-10-01
- Last updated by the study team
- 2026-06-25
Who can join
Age: 18 and older. Sex: any. Healthy volunteers: not accepted.
You may qualify if…
- Adult patients (aged 18 years and older) presenting to the emergency department.
- Complete electronic health record available in the hospital information system (HBYS) containing a detailed anamnesis note with chief complaint, symptom duration, associated symptoms, and relevant medical history.
- A definitive primary diagnosis recorded by the treating emergency physician using ICD-10 codes at the time of patient file closure.
You may not qualify if…
- Emergency department anamnesis notes containing fewer than 50 words or completely lacking substantive clinical content[cite: 1].
- Pediatric cases (age under 18 years)[cite: 1].
- Patients critically ill and triaged to high-acuity resuscitation areas (Emergency Severity Index [ESI] level 1)[cite: 1].
- Clinical notes containing residual identifying information that cannot be fully de-identified, preventing compliance with data privacy regulations[cite: 1].
- Non-independent clinical notes consisting solely of a brief cross-reference to a prior hospital visit without a new history entry[cite: 1].
Where it is running
- Marmara University Pendik Training and Research Hospital — Istanbul, Istanbul, Turkey (Türkiye) (enrolling)
Full record on ClinicalTrials.gov
Trial information comes from ClinicalTrials.gov and is refreshed daily. TrialsForMe does not provide medical care and does not run the studies it lists.