Sleep2 logo Sleep2 logo

Sleep Tracker Accuracy: Why the Criticism is Justified

Manuel Schabus | 21.09.2026

sensorueberblick, uhr, ring, brust, armgurt

A systematic review in Sleep Medicine (Srivali, Duke University, and Cheungpasitporn, Mayo Clinic, 2026) concludes that wrist trackers only explain 2.5 to 16.2 percent of the variance in subjectively experienced sleep quality. In parallel, a class action lawsuit was filed against Oura in the USA in August 2026, alleging misleading accuracy promises. The question of sleep tracker accuracy has thus moved from the specialist journal to the public. This is generally okay.

As a sleep researcher, I say it right away: the criticism is fundamentally justified. The only thing wrong is the conclusion many draw from it. From "most wearables do not reliably measure sleep" it quickly becomes "wearables cannot measure sleep." The difference can be substantiated with data.

What the Sleep Medicine Review Actually Shows

The authors analyzed five studies with around 2,000 participants, four of them with Fitbit models, one with the Apple Watch. Oura or other rings were not included in the review, which is important for the context of the lawsuit. I consider three findings central. First: Tracker data and self-reported sleep quality are only weakly correlated. Second: In one of the included studies, the agreement of a Fitbit with lab measurement (polysomnography) drops from 82 percent in good sleepers to 39 percent in people with insomnia. Third: In this group, total sleep time is overestimated by about 30 minutes and sleep efficiency by about 8 percent.

The authors' conclusion is accordingly cautious: Device data should not be considered a sufficient substitute for validated subjective instruments for assessing sleep quality. For healthy people, the devices at least provide reasonable feedback on sleep duration and regularity. I agree with that.

Why the Criticism is Justified

Those who measure sleep at the wrist primarily measure movement and, depending on the device, an optical pulse signal. Both are only indirectly linked to sleep stages. The pattern has been the same for years: The total sleep time is accurate to about 15 to 30 minutes on the best devices, the rough distinction between sleep and wake works, but short wake phases are underestimated by more than 50 percent, and the sleep stages are close to chance for many devices. What sleep trackers can reliably do today and what they cannot has been described in detail elsewhere.

The fact that it becomes difficult specifically with insomnia is no coincidence. People with chronic insomnia, about 10 percent of adults, sometimes lie awake quietly for a long time. A motion sensor interprets this rest as sleep. The exact group that would most need accurate measurement gets the least accurate values. There is a second problem that a review in Sensors (Lepley et al., 2026, University of Michigan and Hope College) names: Readiness and recovery scores are based on proprietary weightings that are generally without published validation. A score of 78 thus says absolutely nothing that could be verified.

Sleep Tracker Accuracy is Not a Feature of the Device Category

Accuracy is not a feature of the device category. It is a feature of the measurement system. Three factors decide: the quality of the sensor signal, the position on the body, and an algorithm trained against the gold standard, i.e., against polysomnography (PSG) in the sleep lab with EEG, EOG, EMG (and EKG). PSG remains the gold standard for diagnosing sleep disorders.

A sensor on the upper arm or a chest strap delivers a much cleaner heart signal than a sensor on the wrist because there are fewer movement artifacts and less blood flow fluctuation. And a model that has learned from thousands of PSG nights how REM sleep, light sleep, deep sleep, and wake phases map in heart rhythm can read much more from this signal than a motion model. How the four sleep phases physiologically differ explains why this is even possible.

What the Salzburg Comparative Study Shows

Our research group at the University of Salzburg tested exactly this (Topalidis et al., 2025, https://osf.io/preprints/psyarxiv/27wun_v1). The design was deliberately realistic: ambulatory polysomnography at home as the gold standard, five consecutive nights, including disturbance conditions like one hour of forced wakefulness and extended bed time. 90 PSG recordings and over 80,000 30-second epochs were included in the evaluation. Sleep² was compared with Polar Verity Sense and Polar H10 sensors, Oura Ring 3, Apple Watch Series 9, Fitbit Charge 6, Garmin Vivoactive and Venu, WHOOP 4, and the Circul+ Ring.

In terms of total sleep time, sleep² was on average about 6 minutes off from PSG with the Verity Sense, and about 12 minutes with the H10. Other devices overestimated sleep time by up to an hour, especially on restless nights. In the four-class recognition of sleep stages, epoch by epoch against PSG, sleep² with the Verity Sense achieved 83.7 percent, the best result in the entire comparison. This 83.7 percent corresponds to about 95 percent of what two human scorers achieve among themselves in the sleep lab, as their agreement is - surprisingly but true - at a maximum of about 88 percent. A wearable now measures almost as accurately as a second human scorer.

Among the ring and wrist trackers, the Oura Ring 3 performed best with 72.5 percent, just ahead of the Apple Watch Series 9 with 72.3 percent; followed by WHOOP 4 with around 69 percent, Fitbit Charge 6 with around 66 percent, Garmin with around 63 percent, and Circul+ with around 56 percent. Oura is thus a good example that even a ring with an optical sensor on the finger can deliver respectable sleep data if a manufacturer invests in validation.

Not as accurately as a heart signal from the upper arm with an algorithm trained against PSG, and the deviations from PSG fluctuated more from night to night with Oura, but for a wearable on the finger, a very decent result.

I mention these numbers not to disparage other manufacturers; many of these devices measure activity, heart rate, or sleep duration well. I mention them because they show: Wearable is not wearable. The range within the category, from about 56 to 83.7 percent, is greater than the gap between the best wearables and the lab.

What Should Become Industry Standard

From my perspective, the current debate leads to three demands that should apply to everyone, including us. First: validation against PSG, not only in healthy young sleepers but also in disturbed nights and people with insomnia. Second: publication of accuracy per sleep stage, epoch by epoch, instead of a single overall number. Third: transparency instead of proprietary sleep scores. A value whose formula no one knows is an opinion and not a measurement.

Why the second demand is so important is shown by deep sleep. In healthy adults, it accounts for about 20 percent of the night, often below 15 percent by age 50. A device that displays a deep sleep value in minutes every morning should be able to prove how reliably it recognizes this stage exactly. A single overall accuracy can mask weak deep sleep recognition because the most frequent stage dominates the overall number. The same applies to the test conditions: Validation in quiet nights of healthy sleepers says little about how a system processes a night with an hour of wakefulness.

A court will decide on the class action lawsuit against Oura, and I do not presume to judge it. The plaintiffs refer to studies that report sleep stage accuracy near chance; in our own data, Oura performs significantly better than most wrist trackers. As a signal to the industry, however, the lawsuit is clear: accuracy promises must be verifiable, with published and verifiable numbers. Sleep² consciously aligns itself with the validated wearables in this debate because we disclose our numbers and let ourselves be measured by them.

Anyone who wants to measure their sleep objectively and with published accuracy can find sleep² in the App Store and on Google Play. 


Portrait Manuel Schabus

Article by

Manuel Schabus