Chatbots Compared on Knee Osteoarthritis Answers
- •PLoS ONE study compared Perplexity, ChatGPT-5 and Gemini on knee osteoarthritis patient queries
- •ChatGPT-5 produced most readable answers, including FRES:45, FKGL:9.24 and SMOG:8.37
- •Perplexity led quality and reliability scores, with DISCERN: 4, JAMA: 2, EQIP: 92.8
Erdem Maraşlı, E. Ozduran and Volkan Hancı published a 2026 PLoS ONE study comparing Perplexity, ChatGPT-5 and Gemini on answers to knee osteoarthritis questions. Knee osteoarthritis accounts for approximately four-fifths of the global osteoarthritis burden, and the researchers tested whether popular AI chatbots produced patient information that was readable, accurate and reliable.
The study selected 8 keywords from the 25 most frequently used English knee osteoarthritis keywords in Google Trends after excluding repetitive, irrelevant or synonymous terms. The most frequently searched terms were “osteoarthritis of knee,” “knee pain,” and “osteoarthritis knee pain.” Researchers asked those terms as questions to 3 AI-based chatbots and measured readability with Coleman-Liau Index, Automated Readability Index and Linsear Write, among other formulas.
All 3 chatbot systems produced text above the Grade 6 threshold, and the readability difference was statistically significant (p < 0.05). ChatGPT-5 produced the most readable content, with FRES:45, GFOG:11.9, FKGL:9.24, CLI:14.03, SMOG:8.37, ARI:11.37 and LW:7.2.
Perplexity scored significantly higher than ChatGPT-5 across all quality and reliability assessments, with median DISCERN: 4, JAMA: 2, GQS: 4 and EQIP: 92.8. Perplexity also outperformed Gemini on the mDISCERN reliability assessment (p = 0.001), while Gemini and ChatGPT-5 showed no statistically significant difference in reliability and quality surveys.
The authors concluded that knee osteoarthritis answers from popular AI chatbots remain difficult for patients to understand and raised concerns about scientific validity and integrity in medical information. The study said future AI-based tools need effective oversight mechanisms to ensure sufficient quality, robustness and appropriate understandability.