Skip to content
WearablesSep 6, 20264 min read

Smartwatch calorie estimates miss by 15–25%, while readiness scores lack validation

An FIU team measured four popular watches against a metabolic cart and found median absolute percentage errors of roughly 15–25%, with error rising alongside body fat. A separate review found 14 readiness scores across 10 brands with little published validation behind them.

ByMobileTech Desk

Two research groups spent this year asking a question most wearable owners never think to ask: which of the numbers on that screen were measured, and which were guessed? The answers point in the same direction. Many figures beyond heart rate are estimates, and the calorie estimates became less accurate as body fat increased.

58 people, four watches and a metabolic cart

The calorie study came out of Florida International University and ran in PLOS One on 29 July 2026. Jason Kostrna and colleagues put 58 Hispanic adults — 31 women, 27 men, mean age 23, mean BMI 29.73 — on a recumbent cycle ergometer for a 10-minute bout alternating two-minute intervals at roughly 64–76% and 77–95% of maximum heart rate, bracketed by five minutes of rest either side.

The reference was a COSMED K5 portable metabolic analyser doing breath-by-breath indirect calorimetry, which is as close to ground truth for energy expenditure as you get outside a whole-room calorimeter. The watches were an Apple Watch Series 8, a Fitbit Sense 2, a Samsung Galaxy Watch 5 and a Garmin Forerunner 955. Recruitment was stratified by BMI and by Fitzpatrick skin type, which is why the study can say something about both.

The errors were large, and not evenly distributed

Mean bias in kilocalories over that short bout: Garmin 68.61, Samsung 56.76, Apple 21.60. Fitbit's headline bias of 3.14 kcal looks like a win until you see it required removing seven implausible outliers — with them in, the bias was 128.6 kcal. After filtering, Fitbit still showed wide variation and occasionally failed to return a calorie estimate. Reporting on the study noted Fitbit sometimes produced implausible readings, including a single calorie for an entire workout, and sometimes produced nothing at all.

Across the set, median absolute percentage errors ranged from roughly 15% to 25% on the study protocol. Kostrna's own framing to reporters was direct: "You can't treat the calories-burned number it gives you as an accurate number, because it's not." He put the consequence in the terms dieters actually care about — "You could be off hundreds of calories each week and easily end up in a calorie surplus when you think you're in a calorie deficit."

The body-fat finding is the one that matters

Body fat percentage was a strong predictor of error across every device (p < .01), with a significant device-by-body-fat interaction (p = .02) and steeper error slopes for Garmin and Samsung. Skin tone showed no significant main effect (p = .89), though the authors flag that the null "should be interpreted with caution, given the limited sample size."

Their proposed mechanism is worth sitting with, because it inverts a widely repeated assumption about PPG sensors: "Adiposity may have an even greater impact on PPG signal quality than skin pigmentation," with their modelling showing up to 60% signal reduction in severe obesity against roughly 15% attributable to melanin. The paper's conclusion is unsparing — "current consumer devices do not yet provide reliable caloric monitoring for individuals or for research; improving accuracy across body types is essential."

Readiness scores have no paper trail at all

The second piece of evidence is a narrative review by Adam S. Lepley, Fiddy Davis, Amanda C. Melvin and Zheng-Yang Zhao of the University of Michigan and Hope College, published in Sensors (2026, 26:4486). They went looking for validation studies behind 14 composite readiness and recovery scores across 10 manufacturers and reported that these scores "generally lack published validation," with theoretical rationale "emphasized over actual evidence."

Because each brand applies its own unpublished weighting to the same underlying signals, two watches fed identical physiology can hand you opposite verdicts about whether to train. The review's summary of what survives scrutiny is useful as a shopping filter. Reasonably validated: resting heart rate, sleep-versus-wake detection, VO2 max estimates for casual exercisers. Poorly validated: readiness composites, calorie burn at 10–20% error, sleep staging at 50–86% accuracy, cuffless blood pressure, and glucose estimation with no FDA authorisation behind it.

As the authors put it, "most of the numbers on a smartwatch screen are not direct readings of what's happening inside the body."

How to use the watch you already own

Neither team argues for throwing the hardware away. Kostrna was explicit: "They may be imperfect, but these devices can still have an important role to play." What both bodies of work support is a change in which numbers you act on.

Heart rate and time-in-zone survive validation better than anything derived from them, so train off those rather than off a calorie total. If you are eating against a target, round the burn figure down rather than banking it, and resist widening a deficit to match a number that may be a fifth too generous. Watch your own trend line across weeks instead of reacting to today's score, since a consistent bias still tracks direction even when the absolute value is wrong. And treat the composite readiness verdict as a mood ring with good graphic design until a manufacturer publishes the validation work.

The Michigan group's own recommendation is the sensible ceiling on all of this: smartwatch data "should support lab tests, doctor visits, and self-reported information, rather than stand in for them."

More from the desk