LLM Accuracy in Lipedema Patient Education Evaluated
- •DeepSeek outperformed Gemini and ChatGPT with a 72% accuracy rate on lipedema inquiries.
- •Study assessed 25 patient-centered questions using three experts and a four-point rating scale.
- •Researchers concluded that LLM accuracy is insufficient for complex clinical domains requiring expert oversight.
A study published in the journal Phlebology on September 23, 2026, evaluated the accuracy and reproducibility of ChatGPT, DeepSeek, and Gemini in responding to patient inquiries about lipedema, a chronic condition frequently misdiagnosed in clinical settings. The research team led by Rabia Sanır assessed each model across 25 standardized questions. To ensure reliability, researchers performed two separate query sessions for each model, with three independent experts scoring the outputs on a four-point scale.
Results indicated that DeepSeek performed best, providing comprehensive and correct answers 72% of the time, followed by Gemini at 64% and ChatGPT at 56%. While models demonstrated high reliability for general information and diagnostic topics, performance dropped in complex areas such as treatment, follow-up, and long-term maintenance. DeepSeek also displayed the highest level of response consistency across sessions, whereas ChatGPT and Gemini showed greater variability, particularly regarding quality-of-life topics. Cohen's kappa coefficients (a statistical measure of inter-rater agreement) confirmed significant agreement among the experts.
The study concludes that while these LLMs serve as useful tools for initial patient education, they lack sufficient reliability for complex medical decision-making. Researchers emphasized that expert oversight remains mandatory when using AI for patient-centered care, noting that reduced accuracy in nuanced clinical domains currently limits the clinical utility of these models for lipedema patients.