The Sleep Score Is Not Sleep: When the Machine Overrules the Body
You wake before the alarm and lie still for a moment, taking inventory.
Your mind feels clear. Your body feels warm and heavy in the pleasant way that follows a long night. The room is quiet. Somewhere beyond the blinds, morning has begun.
Then you reach for your phone.
The ring says 61.
Suddenly the night changes. What felt restorative becomes deficient. The calm in your body acquires an asterisk. Perhaps the deep sleep was too short. Perhaps the recovery score means the day should be smaller. Perhaps the feeling of being rested was only ignorance waiting for data.
This is the strange authority of the sleep score. It does not merely describe the previous night. It can rewrite it.
Consumer wearables have given millions of people a view into physiology that once required a laboratory. A ring or watch can estimate when sleep began, how long it lasted, how often the body stirred, and how pulse, temperature, breathing, and blood oxygen changed through the night. Across weeks, those estimates can reveal useful patterns.
But a machine that detects sleep is not necessarily a machine that knows whether sleep was good. That difference, obvious once stated, has been blurred by years of dashboards that compress a complicated biological and psychological experience into a number.
A new systematic review from researchers affiliated with Duke University and Mayo Clinic exposes the size of that gap. Its finding is not that sleep trackers measure nothing. It is more unsettling: they may measure real things while missing the thing people most want the number to mean.
What a wrist can see
Inside a sleep laboratory, polysomnography records signals from the brain, eyes, muscles, heart, airflow, breathing effort, and blood oxygen. Trained specialists use those signals to score sleep and identify abnormalities.
A consumer watch or ring has a narrower view. Most infer sleep primarily from movement and optical pulse signals, sometimes supplemented by temperature or oxygen measurements. Proprietary algorithms then translate those inputs into labels such as awake, light, deep, and REM sleep.
This is a feat of engineering. It is also an inference problem.
When the body is moving, wakefulness is relatively easy to recognize. When someone with insomnia lies perfectly still, awake and frustrated in the dark, the same stillness can look like sleep. The device is not being deceitful. It is answering from the evidence available to it.
That helps explain a recurring pattern in validation studies: consumer devices can be sensitive to sleep while remaining much less specific for wakefulness.
In a 2025 laboratory study, 62 adults wore combinations of the Fitbit Charge 5, Fitbit Sense, Withings ScanWatch, Garmin Vivosmart 4, WHOOP 4.0, and Apple Watch Series 8 while undergoing polysomnography. Every device detected more than 90 percent of sleep epochs. Wake specificity, however, ranged from only 29.39 percent to 52.15 percent. Agreement across four sleep stages was fair to moderate.
The headline “more than 90 percent” sounds reassuring until the question changes. A tracker may be good at recognizing that sleep occurred while remaining much less reliable at identifying quiet wakefulness or deciding exactly which stage occupied a given interval.
A separate 2025 meta-analysis pooled 24 studies involving 798 participants. Across devices, researchers found significant differences from polysomnography in total sleep time, sleep efficiency, sleep latency, and wake after sleep onset. The studies were highly heterogeneous, which means no pooled difference should be mistaken for the fixed error of a particular current device. Algorithms change. Hardware changes. Populations change. Performance changes with them.
The durable lesson is not that all wearables are equally inaccurate. It is that “accuracy” is not one thing.
The experience no sensor can directly record
Sleep quality is partly physiological, but it is also experiential. It includes how difficult it was to fall asleep, how often the night felt interrupted, whether sleep seemed restorative, and how a person functions the next day. Clinical insomnia is not diagnosed by finding the wrong percentage of REM sleep on a wrist.
Featured Partner
Invest in the Infrastructure Behind Modern Medicine
As healthcare expands beyond hospital walls, the buildings and campuses supporting that shift are generating compelling returns for investors who move early. The Healthcare Real Estate Fund offers qualified investors direct access to a curated portfolio of medical office, outpatient, and specialty care facilities.
Learn More →The 2026 Duke–Mayo systematic review examined how consumer wearable metrics aligned with validated reports of sleep quality. Five observational studies involving 2,006 adults met the criteria.
Wearable measures explained only 2.5 percent to 16.2 percent of the variation in subjective sleep ratings. Total sleep time had a moderate relationship with same-day diaries, but device measurements did not capture scores from the Pittsburgh Sleep Quality Index well. Reported concordance was 82.4 percent among good sleepers and 39.4 percent among people with insomnia.
The devices also tended to overestimate sleep efficiency, underestimate periods of wakefulness after sleep began, and correlate poorly with sleep-onset latency. The authors rated the certainty of the evidence from low to moderate, an important limitation. Five studies cannot settle the performance of every wearable or every future algorithm.
What they can do is clarify the category error. A sleep tracker measures signals associated with sleep. A person experiences sleep. Those are overlapping domains, not interchangeable ones.
This does not make subjective experience infallible. People can misjudge how long they slept, and sleep-state misperception is a recognized clinical phenomenon. Nor is polysomnography a nightly oracle of how restored someone should feel. Each perspective answers a different question.
The problem begins when a single composite score pretends to answer all of them.
When the number enters the nervous system
Data feels passive, but feedback is an intervention.
In 2014, researchers Christina Draganich and Kristi Erdal told participants that a sensor had measured the quality of their REM sleep. The feedback was fabricated. People told they had slept below average performed worse on selected cognitive tasks than those told they had slept above average. Their belief about the night influenced what happened next.
A 2018 experiment moved closer to the modern wearable experience. Sixty-three people with insomnia received sham positive or negative feedback about their sleep efficiency through an actigraphy-style device. Negative feedback increased reports of sleepiness and fatigue and reduced reported mood and alertness relative to positive feedback. The effects were clearest in subjective daytime symptoms, not across every objective performance measure.
The distinction matters. These studies do not prove that sleep wearables generally damage cognition or cause insomnia. They show that supposedly objective sleep feedback can change how people interpret their condition and experience the following day.
For a healthy sleeper, a low score may be a curiosity. For someone already monitoring fatigue, fearing another bad night, or trying desperately to perfect sleep, it can become evidence for a threat narrative.
Sleep is unusually vulnerable to this loop because it cannot be forced. The harder a person tries to manufacture unconsciousness, the more alert the system can become. Monitoring creates evaluation. Evaluation creates effort. Effort creates arousal. A technology designed to reduce uncertainty can feed the vigilance that keeps sleep away.
The rise of orthosomnia
Sleep clinicians began describing this pattern before modern rings became mainstream.
In a 2017 report, researchers presented three patients who had become preoccupied with improving their wearable sleep data. They called the pattern “orthosomnia,” combining the idea of correct or perfected sleep with the perfectionism seen in orthorexia.
One patient pursued eight hours of sleep according to his tracker and extended his time in bed to make the number rise. That strategy can backfire in insomnia treatment, where excessive time in bed may weaken sleep efficiency and deepen frustration. The clinicians sometimes had to persuade patients that the device was measuring movement, not brain waves, and that lying still with a phone could be recorded as light sleep.
Orthosomnia is not a formal diagnosis, and three cases do not establish its prevalence. The term is useful because it names a recognizable inversion: the pursuit of ideal sleep data begins to interfere with sleep itself.
The most dangerous version is not simply worrying over a low score. It is allowing the score to overrule symptoms in either direction. A reassuring dashboard could delay help for persistent insomnia, loud snoring, witnessed breathing pauses, excessive daytime sleepiness, or another concern. A discouraging dashboard could convince a healthy sleeper that ordinary biological variation is evidence of dysfunction.
The American Academy of Sleep Medicine advises that consumer sleep technologies are not substitutes for medical evaluation. If a product is intended to diagnose or treat a sleep disorder, it should be rigorously tested and cleared for that purpose. Consumer data may still enrich a conversation with a clinician, but it should enter as context, not verdict.
The tracker is most useful when it becomes less important
There is a better relationship with sleep data.
Use it to look for patterns across weeks, not to pass judgment every morning. Bedtime consistency, wake time, approximate duration, resting heart rate, and changes from a personal baseline can be useful signals. They can help test ordinary questions: Does late alcohol fragment the night? Does travel shift the schedule? Does an earlier wind-down change time in bed? Does a persistent pattern match daytime symptoms?
Be more skeptical of fine-grained stage minutes and any composite score that hides its weighting. A label such as “deep sleep” is an algorithmic classification, not a direct reading of the brain. A recovery score may combine several reasonably measured signals while still reflecting design choices about what counts as recovery.
Most importantly, keep the hierarchy intact:
- Persistent symptoms and safety concerns matter even when the score looks good.
- A clinician and validated testing matter when diagnosis is the question.
- Long-term trends are generally more informative than one night.
- A wearable estimate is evidence, not identity.
- Feeling restored is meaningful information, not an error to be corrected by an app.
If checking the score reliably produces anxiety, changes the shape of the day, or encourages more time awake in bed trying to force a better number, taking a break is not technological defeat. It is an experiment in removing the measurement from the system it may be disturbing.
A Diligence Dx check for sleep claims
Before trusting an accuracy claim, ask what was actually tested:
- The target: sleep versus wake, total sleep time, stages, subjective quality, or a disorder?
- The comparator: polysomnography, actigraphy, a diary, a questionnaire, or another wearable?
- The population: healthy adults, athletes, older adults, or people with insomnia or another clinical condition?
- The generation: which device and which algorithm version?
- The metric: sensitivity, specificity, agreement, correlation, mean error, or a marketing composite?
- The independence: who funded the study, who designed it, and what conflicts were disclosed?
- The consequence: is the result useful for observing a trend, or is it being used to make a medical decision?
This is the beginning of the Diligence Dx trust layer for wearables: not a single score declaring a device good or bad, but a visible account of what a claim means, what supports it, and where its authority ends.
The sleep score is not sleep. It is a model of signals produced while sleep may or may not have been happening. Sometimes that model is useful. Sometimes it is wrong. Sometimes the most consequential event occurs after the measurement, when the number reaches the waking mind and begins telling the body how it ought to feel.
The machine deserves attention. It does not deserve the final word.
Practical questions
Are sleep trackers accurate?
They are generally better at estimating whether sleep occurred and tracking broad timing patterns than at detecting wakefulness, classifying every sleep stage, measuring subjective sleep quality, or diagnosing a disorder. Accuracy varies by device, algorithm, population, and metric.
Can a wearable diagnose insomnia?
No. Insomnia involves persistent sleep difficulty, adequate opportunity for sleep, and daytime consequences. Consumer tracker data does not replace a clinical evaluation.
Why can a tracker show sleep when I know I was awake?
Someone lying quietly awake may generate movement and pulse patterns that an algorithm interprets as sleep. This is particularly relevant for people with insomnia.
Should I ignore my sleep stages?
Treat stage estimates as approximate trends rather than exact nightly measurements. Polysomnography uses brain, eye, muscle, breathing, and other signals that wrist wearables do not directly capture.
When should I stop tracking sleep?
Consider a break if checking the data repeatedly increases anxiety, causes you to extend time in bed solely to improve a score, or makes you distrust how you feel. Persistent insomnia, severe sleepiness, breathing pauses, or safety concerns deserve professional evaluation regardless of tracker results.
Sources
- Srivali N, Cheungpasitporn W. “Concordance of wearable device sleep metrics with patient-reported sleep quality: A systematic review.” Sleep Medicine. 2026;144:108941. https://pubmed.ncbi.nlm.nih.gov/41946254/
- Lee YJ, Lee JY, Cho JH, Kang YJ, Choi JH. “Performance of consumer wrist-worn sleep tracking devices compared to polysomnography: a meta-analysis.” Journal of Clinical Sleep Medicine. 2025;21(3):573–582. https://pmc.ncbi.nlm.nih.gov/articles/PMC11874098/
- Schyvens AM, et al. “A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography.” Sleep Advances. 2025;6(2):zpaf021. https://pubmed.ncbi.nlm.nih.gov/40303381/
- Robbins R, et al. “Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults.” Sensors. 2024;24(20):6532. https://pubmed.ncbi.nlm.nih.gov/39460013/
- Baron KG, Abbott S, Jao N, Manalo N, Mullen R. “Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?” Journal of Clinical Sleep Medicine. 2017;13(2):351–354. https://pmc.ncbi.nlm.nih.gov/articles/PMC5263088/
- Draganich C, Erdal K. “Placebo sleep affects cognitive functioning.” Journal of Experimental Psychology: Learning, Memory, and Cognition. 2014;40(3):857–864. https://pubmed.ncbi.nlm.nih.gov/24417326/
- Gavriloff D, et al. “Sham sleep feedback delivered via actigraphy biases daytime symptom reports in people with insomnia.” Journal of Sleep Research. 2018;27(6):e12774. https://doi.org/10.1111/jsr.12774
- American Academy of Sleep Medicine. “Consumer Sleep Technology: An American Academy of Sleep Medicine Position Statement.” https://aasm.org/advocacy/position-statements/consumer-sleep-technology
Featured image credits: Mayo Clinic building and signage by Tony Webster, CC BY 2.0 via Wikimedia Commons; Duke University entrance and signage by Carol M. Highsmith, Library of Congress, no known restrictions on publication.
