Skip to content

Guides · Sleep and recovery

Can Sleep Tracking Make Your Sleep Worse?

Published · 7 min read

Yes, for some people, and the mechanism is not the one most people assume.

The usual worry is accuracy: the ring says you got 42 minutes of deep sleep, the ring is probably wrong, and you have been upset by a number that was never real. That is a genuine problem and it is covered separately in how accurate are wearable sleep stages.

This guide is about something else, and it is stranger. There is experimental evidence that a sleep number changes how you perform even when the number is entirely invented. Not because it was inaccurate. Because you read it.

The experiment that handed people a made-up number

The cleanest evidence comes from Draganich and Erdal, Placebo Sleep Affects Cognitive Functioning, published in the Journal of Experimental Psychology: Learning, Memory, and Cognition in 2014 (PubMed).

164 participants reported how well they had slept the night before. They were then wired up, told they were being measured, and given a figure for their own sleep quality.

What they were told How it was framed
28.7 per cent of the night in REM Above average
16.2 per cent of the night in REM Below average

The figures were assigned at random. They were not measured off anybody. A participant who had genuinely slept badly had the same chance of being handed the good number as anyone else.

Then everyone sat the same cognitive tests.

What moved, and what did not

The group handed the above-average figure scored higher on two timed tasks: the Paced Auditory Serial Addition Test and the Controlled Oral Word Association Task. Both are measures of working memory and verbal fluency under time pressure.

The finding that matters more is the one about the other variable. What participants themselves said about their sleep predicted neither result. The invented number predicted performance. The person's own account of their own night did not.

The study also carried controls, and reporting them honestly is the difference between evidence and a headline. A Digit Span task did not move, which the authors expected. A Symbol Digit Modalities Test also did not move, which they did not expect and say so. So this is a real effect on some measures and not a blanket effect on cognition.

Sleep clinics had already noticed

Three years later, a case series in the Journal of Clinical Sleep Medicine gave the clinical version of this a name.

Baron and colleagues, Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?, 2017 (DOI), described patients arriving at a sleep clinic with self-diagnosed sleep problems built on tracker output. Stretches the device had labelled light or restless were being read as insufficient sleep, and the inferred link between that data and daytime tiredness was turning into what the authors call a perfectionistic quest for ideal sleep.

Be careful how much weight you put on this one. It is three patients. A case series names a pattern, it does not establish how common the pattern is, and the authors were describing something they were seeing rather than counting it.

It is now something that can be measured

The reason to take the 2017 paper more seriously than a three-person case series usually deserves is what happened to it afterwards. In 2025 a group including Kelly Baron, the lead author of the original, published a validated instrument for measuring the thing she had named.

The Bergen Orthosomnia Scale appeared in Frontiers in Sleep (DOI). It was built in three stages: an item pool assessed by 34 sleep experts using the Delphi method, then administered to 994 survey respondents with a mean age of 42, split in half for exploratory and confirmatory factor analysis, then validated in a further sample against instruments for sleep behaviour, personality, obsessive-compulsive traits and health anxiety.

Two factors came out, six items each, making a 12-item scale:

Factor What it captures
Interference The tracking getting in the way of the sleeping
Rigidity Inflexibility about hitting the numbers

Note what those two are. Neither is about the device being wrong. Both are about the relationship with the readout.

The part that should change how you read your own numbers

Put the three papers next to each other and a single conclusion falls out that none of them states on its own.

A sleep score is not a neutral readout of last night. It is an input to the day you then have.

That is not an argument for throwing the ring in a drawer. The data is real and it is useful. It is an argument about what kind of object a score is, and about one specific move that causes most of the trouble: treating a number as a grade.

A grade needs a class. When an app tells you your sleep was below average, the average is somebody else's, assembled from a population you were never in. That is the same structural problem behind why Oura and WHOOP give different scores and behind feeling tired with a good recovery score. The number is being asked to do a job it cannot do.

What your own data can actually tell you here

The useful question is not "was that good" but "was that different, for me, and did anything follow from it".

That reframing survives everything above, because it does not require the absolute number to be accurate or the population comparison to be meaningful. It only requires the device to be consistently wrong in the same direction, which is a far weaker assumption and usually a safe one. A tracker that systematically undercounts your deep sleep by 20 minutes still tells you truthfully when your deep sleep dropped by 40 more than usual.

Three practical consequences:

  1. Read changes against your own history, not scores against a population. This is what a personal health baseline is for.
  2. Give a change time to be a change. One night is noise. The interesting object is the run.
  3. Let "nothing happened" be an answer. If you started something and your own numbers did not move, that is a result about you, and it is more useful than a score that goes up and down for reasons nobody can name. The same logic applies to working out whether a supplement is doing anything.

What this evidence does not establish

Stating the limits plainly, because the papers do.

  • The placebo sleep study was one session with 164 people, and it moved two of four cognitive measures. It does not show that sleep scores harm sleep. It shows that a told number changes near-term performance.
  • The orthosomnia paper is three patients. It names a pattern and describes it well. It does not tell you how common it is, or who is at risk.
  • The Bergen scale is an instrument, not a prevalence study. It gives researchers a way to measure orthosomnia. It does not by itself say how many tracker users have it.
  • None of this says consumer sleep tracking is bad, or that any specific device is inaccurate. That is a separate question with separate evidence.
  • Nothing here is medical advice. Persistent trouble sleeping is worth taking to a clinician, and a tracker is not a diagnosis.

The short version

Handing 164 people a randomly assigned figure for their own sleep changed how they performed afterwards, while their own sense of how they had slept changed nothing. Sleep clinics named the version of this that walks through their door, and there is now a validated 12-item scale for measuring it, built on two factors that are both about the reader rather than the device.

The score is not a readout. It is an input. The way to stop it acting on you is to stop reading it as a grade against strangers and start reading it as a change against yourself.

That is the whole of the design argument behind NUVARD. It reads a change against your own history rather than handing you a grade, and when something has no measurable effect on you it says so, because "this did nothing for you" is a finding and not a gap.

Sources

  • Draganich C, Erdal K. Placebo sleep affects cognitive functioning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 2014. PubMed 24417326
  • Baron KG, Abbott S, Jao N, et al. Orthosomnia: Are Some Patients Taking the Quantified Self Too Far? Journal of Clinical Sleep Medicine, 2017;13(2):351-354. DOI 10.5664/jcsm.6472
  • Guldbrandsen BV, Baron KG, Vedaa O, Bjorvatn B, Pallesen S. Development of a scale for measuring orthosomnia: the Bergen Orthosomnia Scale (BOS). Frontiers in Sleep, 2025. DOI 10.3389/frsle.2025.1640355

Join the waitlist

NUVARD provides wellness intelligence. It does not diagnose, treat, or replace medical care.

← All guides