Why Wearable Sleep Scores Are Lying to Your Clients (And What Actually Predicts Recovery)
The number on the screen is a guess. The way your client feels is data.
Your client texts you at 7:12 a.m.: "Oura said 62. Recovery red. Skipping today?" You glance at their training week, their last three sessions, their stress load, and the fact that they slept seven and a half hours. Everything looks fine. But the ring said 62. And now you're negotiating with a number neither of you can see inside.
This conversation is happening in every coaching practice in 2026. The sleep wearable has become the loudest voice in the room — sometimes louder than the coach, often louder than the client's own body. The problem isn't that wearables are useless. They're genuinely impressive engineering. The problem is what they're being asked to do, and the gap between the marketing claim and the validation literature is wider than most coaches realize.
What the 2024–2025 Validation Studies Actually Found
Robbins et al. (2024), publishing a head-to-head validation of three commercial wearables against polysomnography — the gold-standard, in-lab sleep measurement — found that for two-stage classification (sleep vs. wake), modern devices like the Oura Ring perform reasonably well. But for staging sleep into REM, deep, and light, sensitivity ranged from roughly 50% to 86% across devices, with precision in similar ranges. In plain English: when your client's ring tells them they got 28 minutes of deep sleep, the actual number could be anywhere from 15 to 50 minutes.
Miller et al. (2022) — still one of the most-cited validation papers in the space — examined six popular wearables and found similar patterns: total sleep time estimates are usable, but the per-stage breakdowns that drive most "recovery scores" carry substantial measurement error. Mogavero et al. (2025), in a narrative review published in Clocks & Sleep (MDPI), summarized the current state bluntly: wearables are valuable for trend tracking and population-level patterns, but interpreting any individual night's stage breakdown as ground truth is not supported by the evidence.
The most important consequence isn't that scores are wrong sometimes. It's that the proprietary "recovery" and "readiness" scores aggregate already-noisy stage estimates with HRV, resting heart rate, and respiratory rate, then run them through an opaque algorithm that hasn't been independently validated. Your client is being told a number that the company computing it can't fully defend.
The Score Isn't the Problem. The Behavior Is.
The clinical harm from wearable sleep tracking isn't usually inaccurate data. It's the behavior change the data drives. Researchers have been describing this for years under the term "orthosomnia" — a preoccupation with achieving perfect sleep scores that paradoxically degrades sleep quality. The pattern is consistent: clients see a low score, feel anxious about being under-recovered, ruminate about sleep the next night, and then sleep worse — confirming the score and reinforcing the loop.
For coaches, the operational harm shows up in three places. Clients skip training they could handle. Clients push training they shouldn't. Clients begin to outsource a coaching decision — "should I train today?" — to a device that doesn't know their goals, their context, or their history with you.
What Actually Predicts Recovery (For a Gen-Pop Client)
The recovery literature, when you strip away the device-specific marketing, points to a much shorter and more boring list of predictors than the apps suggest. Total sleep duration is the strongest single signal. Sleep regularity — going to bed and waking up at consistent times — has accumulated enough evidence in the last five years to rival duration as a predictor of next-day function. Subjective sleep quality (a one-question rating from the client) tracks performance and mood outcomes nearly as well as instrumented metrics in non-clinical populations. And HRV trend over weeks, not single-day readings, is where the actual coaching signal lives.
None of those four require a $400 ring. They require asking the client.
The Three-Question Morning Check-In That Outperforms a Wearable
For most gen-pop clients, three questions answered honestly at the start of the day will give you more useful information than any score:
1. "On a scale of 1–10, how rested do you feel?" Subjective ratings track training tolerance with surprising fidelity. A 7 means train as planned. A 5 means dial back load by ~10%. A 3 means restorative work only.
2. "Did you fall asleep within 20 minutes and stay asleep?" This single question captures most of what stage-based scores try to estimate, without the noise.
3. "What's your stress load today outside the gym?" A high-deadline workday compounds with training stress in ways no wearable detects.
This isn't anti-tech. It's pro-signal. The wearable can stay on. It just doesn't get to be the deciding vote.
How to Coach a Client Who Already Trusts the Score
You can't talk a client out of a number they've internalized. But you can re-rank it. Three moves work:
Reframe the score as one input among many. "The ring is one data point. Your sleep duration, your subjective rating, and how the last session felt are three more. We don't make decisions on one data point."
Track score vs. actual session quality for four weeks. Have the client log their recovery score and then rate the session 1–10 afterward. Within a month, most clients see for themselves that the correlation is weaker than they expected. The score loses its grip when the client gathers their own counterevidence.
Set a "score floor" instead of a daily veto. Agree in advance: if score is below, say, 40 for three consecutive days and subjective rating is below 5, that triggers a deload conversation. Otherwise the score is information, not instruction.
What to Tell Clients About Their Tracker
You don't need to be the coach who hates wearables. You need to be the coach who knows what they're for. The honest, evidence-led story is: trackers are good at trend tracking — sleep duration over weeks, HRV over months, identifying nights that are dramatically outside the client's normal range. They are not good at telling your client whether to train today.
The validation literature (Robbins 2024, Miller 2022, Mogavero 2025) supports tracking. It does not support outsourcing daily decisions to the score. That distinction is the entire coaching conversation.
The Real Recovery Skill
The clients who recover well aren't the ones with the best wearables. They're the ones with the most consistent sleep schedules, the lowest stress reactivity, and a coach who knows the difference between a noisy data point and a meaningful signal. Build that skill in your clients and the score becomes background. Let the score build that skill for them and you become a customer-service rep for an algorithm.
Your job is to coach the human. The ring is just along for the ride.
Selected References
Robbins R, et al. (2024). Accuracy of three commercial wearable devices for sleep tracking in healthy adults. Sensors. PMC11511193.
Miller DJ, et al. (2022). A validation of six wearable devices for estimating sleep, heart rate and heart rate variability. Sensors. PMC9412437.
Mogavero MP, et al. (2025). Beyond the sleep lab: a narrative review of wearable sleep monitoring. Clocks & Sleep (MDPI).
Connect with Coach Camp
Coach Camp is where evidence-led coaches sharpen their craft and grow sustainable practices. If this resonated, plug in:
🌐 Web: coachcamp.net
🎥 YouTube: youtube.com/@coachcamphq
🐦 X: x.com/coachcamphq
📸 Instagram: instagram.com/coachcamp.app
🏫 Skool community (free): skool.com/coachcamp