LLMs Tested On Sprint Swimming Plans
- •Four LLM tools generated race-day plans for a simulated 100-m freestyle swimmer
- •32 swimming experts rated anonymized warm-up and cool-down plans using a 40-item instrument
- •ChatGPT had highest quality profile, while Doubao produced more stable but lower-rated outputs
Researchers Tian-Yue Zhao, Yi-Yao Peng, Zi-Xin Chen and co-authors reported in Frontiers in Physiology on 2026-08-27 that four publicly accessible LLM-based AI tools produced uneven race-day warm-up and immediate post-race cool-down plans for a simulated 100-m freestyle swimmer. The study evaluated whether ChatGPT, Gemini, DeepSeek and Doubao could generate plans with physiological relevance, content quality, safety and repeated-generation stability for competitive swimming preparation.
The researchers used a simulation-based content-analysis design with a within-rater repeated-measures structure. ChatGPT, Gemini, DeepSeek and Doubao each generated five independent outputs from the same standardized detailed prompt, and each output contained both a warm-up plan and an immediate post-race cool-down plan. After anonymization and randomization, 32 experts with coaching or teaching experience in 100-m sprint swimming rated the plans with a 40-item multidimensional rating instrument.
The analysis assessed inter-rater reliability with intraclass correlation coefficients and Krippendorff’s alpha, then used a cumulative link mixed model (ordinal outcome regression) with fixed effects for AI model, section and their interaction, plus random intercepts for rater, item and generation ID. The model found a significant model × section interaction (p <.001), meaning differences among AI tools changed between warm-up and cool-down sections. ChatGPT had the highest descriptive overall quality, scoring higher than DeepSeek and Doubao for warm-up and higher than all three other models for cool-down.
Exploratory domain-level analyses found differences across all 10 content-quality domains. Generation-level stability did not match expert-rated quality: Doubao was the most stable model, while ChatGPT had the most favorable overall quality profile but greater repeated-generation variability. The authors concluded that AI-generated swimming race-day plans should not be used as stand-alone prescriptions and should remain subject to qualified coach supervision for intensity, sequencing, recovery, risk management and individual adaptation.