How Accurate Are Wearable Sleep Stages?
Published · 5 min read
Reasonably good at telling sleep from wake, and much weaker at telling one stage from another.
That gap is the whole answer, and it is the reason two devices on the same body on the same night can report deep sleep figures that differ by an hour. Neither of them is broken. They are estimating something they cannot observe directly, from signals that only partly carry it, and they are using different rules to do it.
Which means the useful question is not whether your deep sleep number is correct. It is which parts of your sleep data are solid enough to act on, and which are best read as a rough shape.
What the device can actually see
A sleep laboratory scores stages from brain activity, eye movement and muscle tone, read from electrodes. Those are the signals that define the stages in the first place.
Your wearable has none of them. It has movement, a pulse read from the skin, the variability between beats, usually skin temperature, sometimes blood oxygen and breathing rate. From those it infers what the brain was probably doing.
That inference works well for the coarse question. Whether you were asleep or awake produces a large, obvious difference in movement and heart rate, and consumer devices agree with laboratory scoring on it most of the time.
It works less well for the fine question. Light sleep, deep sleep and REM are distinguished in the laboratory by brain patterns, and their peripheral signatures overlap considerably. REM has its own heart rate signature and is the stage wearables generally do best on after wake. Deep sleep is the hardest, and it is also the number people care most about, which is an unfortunate pairing.
Why two devices disagree so completely
Three separate reasons, and they compound.
Different sensors. A ring reads from a finger, a watch from a wrist, a chest strap from the torso. Signal quality, contact and motion artefact are different at each site, so the raw material differs before any interpretation starts.
Different algorithms. Every manufacturer has its own model, trained on its own data, tuned against its own reference. There is no shared standard that says what counts as deep sleep from a wrist, so each company answers the question slightly differently and none of them is obliged to agree.
Different definitions of the night. Devices disagree about when sleep began, whether a quiet stretch on the sofa counted, and how to treat brief awakenings. A twenty minute difference in where the night starts moves every stage total inside it.
None of this makes the data useless. It makes cross device comparison meaningless, which is a different problem with a simple answer: do not do it. Our guide on why Oura and WHOOP give different scores covers how far that disagreement goes across metrics generally.
What the numbers are good for
Stage data is weak as an absolute figure and considerably stronger as a trend on one device.
The absolute number answers a question it cannot really answer. Whether you got seventy two minutes of deep sleep last night is a claim about brain activity that was never measured, and the error bars around it are wide enough that the figure should not be the thing you act on.
The trend answers a question it genuinely can. Whether this week ran lower than your own typical fortnight, on the same device worn the same way, is a comparison between like and like. Whatever bias the algorithm carries is present on both sides of that comparison, so it largely cancels out. That is why a personal baseline is more informative than any single night, and why comparing your figure to a published population range is close to meaningless. Our guide on what a personal health baseline is covers how that comparison is built.
Sleep timing and duration are the sturdiest things your device reports, and they are also the parts most people skip past on the way to the stage breakdown.
How to read your own sleep data without over reading it
Start with what is reliable and work outward.
Total sleep and timing first. When you went to sleep, when you woke, how long you were down, and how consistent that is across a week. This is measured rather than inferred, and it is the part with the strongest relationship to how you actually function.
Then continuity. How fragmented the night was. Also comparatively well captured, because awakenings show up clearly in movement and heart rate.
Then the stage shape, as a pattern rather than a score. Whether your deep sleep is concentrated early, whether REM is showing up in the second half, whether the distribution has shifted from your own normal. Shape and change are more trustworthy than the totals they are made of.
Never across devices, and never against a population chart. A stage figure is only comparable to other figures produced by the same algorithm on the same body.
Where this leaves you
Your device is not lying to you and it is not measuring your brain. It is producing a well informed estimate from the outside, and that estimate is dependable at the coarse scale, indicative at the fine one, and only interpretable against your own history.
The practical failure is not inaccuracy. It is that a single night arrives on a screen as a verdict, with no indication of which parts of it are solid, and no comparison against the only reference that makes it meaningful, which is you.
NUVARD is built on that reference. It learns what is normal for you across every signal your devices already produce, so a night is read against your own baseline rather than against a population range, and a change is reported when it is a change rather than when a number happened to move.
NUVARD is coming soon.