Oura’s director of health science published a personal defence of consumer sleep wearables on the company’s blog 3 days after a proposed class action was filed against Oura on August 20, challenging the company’s sleep-stage accuracy figures.
For many years, this site has pointed out that sleep staging used by most wearable companies cannot be lab-grade, and has pushed readers to understand that even the top-rated lab estimates with human scorers only agree about 83% of the time, meaning wearables are working against a hard ceiling before they even start.
Four figures set out what the science actually says, and where Oura’s own numbers sit within it.
- 79% is Oura’s published figure for sleep-stage accuracy when its device is compared with a sleep lab across four categories: awake, light sleep, deep sleep and REM, that’s 79% accuracy against a human benchmark of about 83%. It comes from Oura’s own studies and is the figure cited in the Surber complaint seeking class-action status in California.
- 83% is the sleep-science benchmark for agreement between two trained sleep-lab technicians scoring the same night. This is a human-scorer figure, not a device figure, and Oura’s own science team cited the identical number this week, arguing that its own 79% four-stage accuracy already represents a high level of performance against that human ceiling.
- Oura’s published accuracy is 92% to 96% when the question is simply whether a person is asleep or awake, rather than which stage of sleep they are in. Oura has also used a 95% figure in its marketing, which is cited in the lawsuit. Asleep-versus-awake is a much easier task than distinguishing sleep stages, so keep that distinction in mind.
- 53.18% is the sleep-stage accuracy reported in an independent clinical study cited in the complaint, run on a clinical population rather than healthy volunteers. Oura did not conduct or fund the study.
The specific weak point for wearables generally, Oura included, is telling quiet wakefulness apart from light sleep. Independent studies of multiple ring and wrist trackers have found this confusion is the most common error pattern across the category, not something unique to any one brand. Oura’s own healthy-population validation studies, by contrast, show strong REM detection.
The bottom line is that no consumer wearable matches the agreement rate that trained human scorers reach with each other in a sleep lab.
Last Updated on 23 August 2026 by the5krunner

tfk is the founder and author of the5krunner, an independent endurance sports technology publication. With 20 years of hands-on testing of GPS watches and wearables, and competing in triathlons at an international age-group level, tfk provides in-depth expert analysis of fitness technology for serious athletes and endurance sport competitors. ID


