Google Launches SensorFM for Wearable Health Data
- •Google introduced SensorFM, a foundation model trained on over one trillion minutes of wearable sensor data.
- •The model generalizes across 35 health tasks, outperforming traditional supervised baselines on 34 of them.
- •An agentic classroom of LLMs automatically builds prediction heads, improving performance on 28 of 35 tasks.
Google researchers unveiled SensorFM, a foundation model (AI trained on vast datasets for broad tasks) designed to learn general-purpose human physiology representations from wearable device data. The model was pre-trained using more than one trillion minutes of sensor signals collected from five million consented participants between September 2024 and September 2025. By utilizing data from over 20 Fitbit and Pixel Watch models across 100 countries and all 50 U.S. states, the research team created the largest and most diverse wearable health dataset to date.
SensorFM ingests 34 aggregate features derived from five sensor modalities, including photoplethysmography (PPG) and accelerometry, to capture data like heart rate, blood-oxygen saturation, sleep stages, and skin temperature. The model employs self-supervised reconstruction via the Adaptive and Inherited Masking (AIM) framework, which treats real-world missing data—a common artifact of wearable usage—as a natural input rather than an error to be imputed or discarded.
Scaling experiments demonstrated that increasing model capacity and data volume leads to predictable improvements, with the largest model variant, SensorFM-B, reducing reconstruction loss by 31% compared to the smallest version. In evaluations across 35 discriminative health tasks, SensorFM-B outperformed feature-engineered baselines on 34 tasks. The embeddings proved particularly effective for identifying health signals in hard-to-measure conditions like depression and anxiety.
To automate the adaptation of these embeddings, the team implemented an agentic "classroom"—a system of competing and collaborating LLM agents that iteratively generate and refine code to build task-specific prediction heads. This system tested over 30,000 candidate solutions and outperformed simple linear probes on 28 of 35 evaluated tasks.
Finally, researchers tested SensorFM as a grounding tool for a Personal Health Agent. In a study involving 31 participant profiles, clinicians blinded to the condition rated summaries generated by the agent. The inclusion of SensorFM predictions significantly improved the quality of health summaries across all five evaluated rubric dimensions—context, relevance, justifiability, personalization, and potential for harm—reaching performance levels statistically comparable to grounding the agent in ground-truth measurements.