RESEARCH ARTICLE
Nobel Med 2026; 22(2): 125-130

EVALUATION OF THE ACCURACY OF CHATGPT-5.2 IN THE DIAGNOSIS AND TREATMENT OF PES PLANUS: AN ARTIFICIAL INTELLIGENCE–BASED STUDY

Muhammed Taha Demir
ABSTRACT
Objective: This study aimed to systematically assess the accuracy of ChatGPT-5.2 (OpenAI, USA) in responding to questions on the diagnosis, evaluation, and treatment of pes planus, and to investigate the potential role of artificial intelligence models in orthopaedic clinical practice.

Material and Method: A total of 30 five-option multiple-choice questions covering the diagnosis, clinical evaluation, radiological findings, and treatment approaches of pes planus were prepared based on current literature. The questions were developed through a multi-stage expert evaluation process by three orthopaedic specialists, one of whom had specific experience in foot and ankle surgery, and were subsequently submitted to the ChatGPT-5.2 model. Assuming a 20% probability of random success for five-option questions, accuracy was analyzed using a binomial test, and the 95% confidence interval was calculated using the Wilson method. Differences in accuracy rates across question categories were analyzed using ffisher’s exact test.

Model responses were independently evaluated by three experts, and interobserver agreement was assessed using ffleiss’ kappa coefficient.

Results: The ChatGPT-5.2 model correctly answered 24 of 30 questions, yielding an overall accuracy of 80% (95% CI: 62.7%–90.5%). Under the assumption of a 20% random success rate, the observed accuracy was found to be significant using a one-sided binomial test (p<0.001). Category-based analysis demonstrated high performance in definition, anatomy, and clinical evaluation questions, whereas lower accuracy was observed in treatment and surgical indication items. Inter-rater agreement among experts was at a near-perfect level (ffleiss’ =0.89). Half of the incorrect responses were classified as high clinical risk, all of which were within the treatment/surgical category. ChatGPT-5.2 performance differed significantly across categories (p=0.048).

Conclusion: ChatGPT-5.2 may serve as a useful supplemental tool for orthopaedic information and patient education regarding pes planus; however, it should not replace specialist clinical judgment.

LIVER AND SYSTEMIC DISEASES

05-09

Sebati Özdemir

REVIEW Nobel Med 2013; 9(2): 5-9

AGING KIDNEY: SENESCENCE OR DISEASE?

10-14

Meltem Gürsu, Rümeyza Kazancıoğlu, Savaş Öztürk

REVIEW Nobel Med 2013; 9(2): 10-14

IMPORTANCE OF HOLOTRANSCOBALAMIN (HOLOTC) MEASUREMENTS IN EARLY DIAGNOSIS OF COBALAMIN DEFICIENCY, ESPECIALLY IN PATIENTS WITH BORDERLINE VITAMIN B12 CONCENTRATIONS

15-20

Faruk Sönmezışık, Esma Sürmen Gür, Burak Asıltaş

RESEARCH ARTICLE Nobel Med 2013; 9(2): 15-20

EVALUATION OF PHYSICAL GROWTH IN PATIENTS WITH FAMILIAL MEDITERRANEAN FEVER

21-25

Celalettin Koşan, Oğuzhan Sepetçigil, Atilla Çayır, Avni Kaya, Behzat Özkan

RESEARCH ARTICLE Nobel Med 2013; 9(2): 21-25

COMPARISON OF RISK INDEXES USED IN DETERMINING THE POSTOPERATIVE RESPIRATORY INSUFFICIENCY RISK

26-31

Gülsüm Kavalcı, Cavidan Arar, Alkın Çolak, Nesrin Turan, Cemil Kavalcı

RESEARCH ARTICLE Nobel Med 2013; 9(2): 26-31

THE EVALUATION OF PRENATAL AND ENVIROMENTAL RISK FACTORS IN CHILDREN WITH ASTHMA, ALLERGIC RHINITIS AND BRONCHIAL ASTHMA ABSTRACT

32-37

Mehmet İbrahim Turan, Müferet Ergüven, Mehmet Özdemir

RESEARCH ARTICLE Nobel Med 2013; 9(2): 32-37

REOPERATIONS AND MORBIDITY IN THYROID SURGERY

38-42

Serkan Teksöz, Murat Özcan, Aytül Sargan, Yusuf Bukey, Recep Özgültekin, Ateş Özyeğin

RESEARCH ARTICLE Nobel Med 2013; 9(2): 38-42

EFFECTS OF RAMADAN FASTING ON BLOOD PRESSURE CONTROL, LIPID PROFILE, BRAIN NATRIURETIC PEPTIDE, RENAL FUNCTIONS AND ELECTROLYTE LEVELS IN HYPERTENSIVE PATIENTS TAKING COMBINATION THERAPY

43-46

İbrahim Faruk Aktürk, İsmail Bıyık, Cüneyt Koşaş, Ahmet Arif Yalçın, Mehmet Ertürk, Fatih Uzun

RESEARCH ARTICLE Nobel Med 2013; 9(2): 43-46

MANAGEMENT OF A LARGE OUTBREAK CAUSED BY NOROVIRUS AND CAMPYLOBACTER JEJUNI OCCURRED IN A RURAL AREA IN TURKEY

47-51

İbak Gönen

RESEARCH ARTICLE Nobel Med 2013; 9(2): 47-51

FREQUENT CD99 AND FLI-1 EXPRESSIONS IN DIFFUSE LARGE B-CELL LYMPHOMA AND THEIR ASSOCIATION WITH PROLIFERATIVE AND APOPTOTIC RATES

52-56

Ufuk Berber, İsmail Yılmaz, Tolga Tuncel, Zafer Küçükodacı Aptullah Haholu

RESEARCH ARTICLE Nobel Med 2013; 9(2): 52-56
  • Pubmed Style
    Muhammed Taha Demir. [PES PLANUS TANI VE TEDAVİSİNDE CHATGPT-5.2’NİN DOĞRULUĞUNUN DEĞERLENDİRİLMESİ: YAPAY ZEKA TABANLI BİR ÇALIŞMA]. Nobel Med 2026; 22(2): 125-130, Turkish.
  • Web Style
    Muhammed Taha Demir. [PES PLANUS TANI VE TEDAVİSİNDE CHATGPT-5.2’NİN DOĞRULUĞUNUN DEĞERLENDİRİLMESİ: YAPAY ZEKA TABANLI BİR ÇALIŞMA]. www.nobelmedicus.com/en/Article.aspx?m=1716 [Access: Mayıs 24, 2021], Turkish.
  • AMA (American Medical Association) Style
    Muhammed Taha Demir. [PES PLANUS TANI VE TEDAVİSİNDE CHATGPT-5.2’NİN DOĞRULUĞUNUN DEĞERLENDİRİLMESİ: YAPAY ZEKA TABANLI BİR ÇALIŞMA]. Nobel Med 2026; 22(2): 125-130, Turkish.
  • Vancouver/ICMJE Style
    Muhammed Taha Demir. [PES PLANUS TANI VE TEDAVİSİNDE CHATGPT-5.2’NİN DOĞRULUĞUNUN DEĞERLENDİRİLMESİ: YAPAY ZEKA TABANLI BİR ÇALIŞMA]. Nobel Med (2026); 22(2): 125-130, [cited Mayıs 24, 2021], Turkish.
  • Harvard Style
    Muhammed Taha Demir. (2026) [PES PLANUS TANI VE TEDAVİSİNDE CHATGPT-5.2’NİN DOĞRULUĞUNUN DEĞERLENDİRİLMESİ: YAPAY ZEKA TABANLI BİR ÇALIŞMA]. Nobel Med, 22(2): 125-130, Turkish.