Skip to content

Guides · Sleep

Why Is My Deep Sleep So Low?

Most likely because deep sleep is the number your device is least able to get right, and because the figure you are comparing yourself against was never meant to apply to you.

Deep sleep is the stage consumer wearables classify least reliably. Two devices worn on the same body, on the same night, can report deep sleep totals that differ by more than an hour, and the direction of the error is a property of the device rather than of your night. Before you treat a low reading as a fact about your body, it is worth knowing how large the measurement error actually is.

That is the short answer. The rest of this is the evidence for it, what a low number does and does not tell you, and what to read instead.

What happens when you check a wearable against the lab

The reference standard for sleep staging is polysomnography, an overnight lab recording of brain activity, eye movement and muscle tone. It is the only method that observes the signals sleep stages are actually defined by. A wearable does not see any of them. It infers stages from movement and from pulse measured at the skin, then applies a manufacturer's algorithm that is not published.

In 2025, a study published in the journal SLEEP Advances put six wrist-worn devices against polysomnography directly. Sixty-two adults each spent a single night in a sleep laboratory wearing two to four of the following: Fitbit Charge 5, Fitbit Sense, Withings Scanwatch, Garmin Vivosmart 4, Whoop 4.0 and Apple Watch Series 8. Every device was scored against the lab recording made simultaneously on the same person.

For deep sleep, measured in minutes against the lab on the same night:

  • Garmin Vivosmart 4 reported 44.44 minutes more deep sleep than the lab found, on average, across 25 participants.
  • Whoop 4.0 reported 31.49 minutes more, across 40 participants.
  • Apple Watch Series 8 reported 25.20 minutes less, across 20 participants.
  • Fitbit Sense was 3.89 minutes low and Fitbit Charge 5 was 2.19 minutes low. Neither difference reached statistical significance.

The Withings Scanwatch is left out of that list on purpose. It does not separate deep sleep from REM at all, so its deep sleep figure is the two stages added together and cannot be compared with the others.

Read the first three lines again. On the same night, in the same room, one device overstates deep sleep by roughly three quarters of an hour while another understates it by close to half an hour. The gap between them is larger than most people's entire deep sleep total.

Why deep sleep specifically is the hard one

Across all six devices the same overall pattern held. Every device was good at answering "is this person asleep", with sensitivity above 90 percent. Every device was much weaker at answering "asleep in which stage", with specificity between 29.39 percent and 52.15 percent. Overall agreement with the lab, measured as Cohen's kappa, ranged from 0.21 to 0.53, which the authors describe as fair to moderate.

The mechanism is straightforward once stated. Deep sleep is a brain-state category. Your device is inferring it from how still you are and from what your pulse is doing. Those correlate with deep sleep, and correlation is enough to produce a plausible number every morning, but it is not the same as observing the thing itself. A still hour is not necessarily a deep hour, and the algorithm deciding which is which is proprietary and differs between manufacturers.

This is the same structural issue that makes two health apps disagree about the same night. It is not that one is broken. It is that each is running a different unpublished algorithm over a different sensor, and the disagreement between them is a fair estimate of how uncertain both are.

What the researchers concluded, and what they did not

Their own conclusion is worth quoting rather than paraphrasing: all the devices tested "can benefit from further improvement for multistate categorization", and the ones with better agreement "could be effectively used to track prolonged and significant changes in sleep architecture".

That sentence contains the whole practical answer. These devices are not reliable for judging one night against a target. They are considerably more useful for noticing that something has shifted across many nights. A single low deep sleep figure sits comfortably inside the measurement error. A sustained change in your own trend does not.

What the study does not say is that any device is useless, or that deep sleep does not matter, or that one brand is correct and the others wrong. Two of the six were within four minutes of the lab on average. The point is not that the numbers are fiction. The point is that the error bars are wide, device-specific, and invisible in the app.

The number you are comparing against is probably not yours

There is a second problem sitting on top of the measurement one, and it is the more common cause of an unnecessary worry.

Most apps present deep sleep against a general adult range. Ranges like that are built from population data, and a population range answers a question about a group. It cannot tell you what is ordinary for you. Deep sleep varies substantially between healthy people, and it declines with age in most of them, so a figure that is unremarkable for one person can sit below the printed range for another who is entirely fine.

The comparison that carries information is you against your own history, which is what a personal baseline is for. Fourteen ordinary nights of your own data tells you more about whether tonight was unusual than any published range will. It also does something a range cannot: it absorbs your device's particular bias. If your watch runs 25 minutes low on everybody, it runs 25 minutes low on you too, consistently, and a comparison against your own history quietly cancels that out. A comparison against a population range does not.

The cost of over-reading it

There is a documented failure mode here and it is worth naming. In 2017 a group of sleep researchers led by Kelly Baron published a case series in the Journal of Clinical Sleep Medicine describing patients who arrived seeking treatment for sleep problems they had identified from their tracker data. They coined the term orthosomnia for it, built from the same root as orthorexia: a preoccupation with achieving the correct sleep.

Orthosomnia is not a formal medical diagnosis and should not be treated as one. It is a description of a pattern the researchers kept seeing, in which the effort to improve a sleep number costs more sleep than the original problem did. Given the error margins above, this is a real risk and not a hypothetical one. Anxiety about a low deep sleep reading is capable of producing a genuinely worse night, at which point the number and the worry begin to feed each other.

What to read instead

Read total sleep time first. It is the figure every device gets closest to right, and it is the one most likely to explain a poor night on its own.

Read your own trend, not the night. Compare this week against your own recent weeks. If deep sleep has been drifting down across a fortnight, that is a signal worth attending to. One low morning is inside the noise.

Check whether anything else moved. A body under real load tends to move several things at once: resting heart rate, heart rate variability, sleep efficiency, temperature. If deep sleep dropped and nothing else did, measurement is the more likely explanation. This is the same corroboration logic that applies when HRV drops but you feel fine.

Do not compare across devices. Given a spread of nearly seventy minutes between devices on the same night, a deep sleep figure from one device and a deep sleep figure from another are not the same measurement and should not be put in the same sentence.

Weigh how you feel at least as heavily. Alertness through the day is a cruder instrument than a sleep lab and a better one than an unvalidated stage estimate.

When a low reading is worth a second look

Understanding the artefact is for seeing past it, not for explaining everything away.

A sustained downward trend in your own deep sleep across weeks, appearing alongside daytime sleepiness, unrefreshing sleep, loud snoring or witnessed pauses in breathing, is worth raising with a doctor. So is any persistent change that is new for you. In that conversation the useful thing to bring is the trend and its timing, not the absolute minutes, because your clinician has no more reason to trust your device's stage algorithm than you do.

A wearable is a reasonable instrument for noticing that a question exists. It is not the instrument that answers it.


Where NUVARD sits on this

NUVARD does not diagnose anything and is not a substitute for medical care. What it does is treat the problem above as the actual problem: a number is not useful until you can see what went into it and what it is being compared against.

That means learning your baseline from your own history rather than scoring you against a population range, reading signals together rather than one card at a time across 300 or more connected devices and apps, and testing what you take against those signals so that NO SIGNAL is a result you are actually told, rather than a finding quietly left out.

Nothing above validates NUVARD's own approach. That is precisely the standard this guide has applied to everyone else's, and it applies here too.

NUVARD launches 21 August. The waitlist gets the download link first: nuvard.ai

Sources

  • Schyvens A M, Peters B, Van Oost N C, Aerts J M, Masci F, Neven A, Dirix H, Wets G, Ross V, Verbraecken J. "A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography." SLEEP Advances, Volume 6, Issue 2, April 2025, zpaf021. https://doi.org/10.1093/sleepadvances/zpaf021
  • Baron K G, Abbott S, Jao N, Manalo N, Mullen R. "Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?" Journal of Clinical Sleep Medicine, Volume 13, Issue 2, 2017, pages 351 to 354. https://doi.org/10.5664/jcsm.6472

Join the waitlist →

NUVARD provides wellness intelligence. It does not diagnose, treat, or replace medical care.

← All guides