LLM Agreement with Anesthesiologist ASA Scoring Evaluated
- •Study compares five LLMs against expert anesthesiologist ASA physical status scoring
- •Copilot achieved highest agreement with 15-year expert (Kappa = 0.82)
- •DeepSeek and Gemini showed significantly lower agreement in 100 simulated cases
Muhammed Emin Zora evaluated the reliability of LLMs in assigning American Society of Anesthesiologists (ASA) physical status scores for preoperative risk assessment. The study utilized 100 simulated patient cases across ASA classes I–V, comparing five AI models—ChatGPT, Copilot, Grok, DeepSeek, and Gemini—against assessments from an anesthesiologist with 15 years of experience.
Results showed significant performance variance (p < 0.001), with Copilot achieving the highest agreement (Cohen’s Kappa = 0.82), followed by ChatGPT and Grok. DeepSeek and Gemini displayed lower performance levels. Discrepancies occurred mainly between adjacent categories, such as ASA II–III and ASA III–IV. The study concludes these models show promise in controlled settings but require further multicenter validation using real-world data before clinical adoption.