Here’s a very easy way to get near 100% sleep vs. wake accuracy from a wristworn sleep tracking algorithm over the course of the night:
Say everything from lights-out to lights-on is sleep
Test only on healthy sleepers
This is a pretty good algorithm because healthy sleepers fall asleep quickly and stay asleep the whole night. You’ll probably get >90% sleep/wake accuracy from this alone. You will miss 100% of the awakenings that happen for them during the night, but these are healthy sleepers, so there aren’t many of those. If they’re awake for five minutes in an 8-hour night, you can miss all five of those minutes and still report 98.958% accuracy.
Of course, you’ll do terribly if you apply this algorithm to a person lying awake in bed for hours, staring at the ceiling and ruminating about every mistake they’ve ever made. But you’re testing only on perfectly healthy sleepers, so that person isn’t there.
The above is why people in the sleep field have moved away from sleep vs. wake accuracy as a metric to brag about: it’s extremely low information. A great algorithm and a terrible one can look the same, depending on who’s in your dataset. You can get very high accuracy with the dumbest algorithm you can think of.
In place of accuracy, people often ask: What fraction of the true wakes did you correctly label as wake? And what fraction of true sleeps did you correctly label as sleep? These numbers are specificity and sensitivity by convention in the sleep world, and our “everything is sleep” algorithm would have a very good sensitivity (gets all true sleep right) and ass specificity (gets all the true wake wrong). An algorithm like “if you are moving, you are awake; otherwise you are asleep” gets around 90% sensitivity and 40% specificity. In general, if you tell me that you’ve made an algorithm for sleep using wristworn signals like acceleration and heart rate, that has 95% sensitivity and a specificity over 65%, I’m impressed.
But this, too, is hackable. All you have to do to get an amazing specificity is include people with a lot of extremely easy-to-classify (high motion) wake in your evaluation population. You can do this by starting the window of time you’re going to use to calculate metrics a bit early (like when participants are walking around getting ready for bed instead of at lights-out) or including people who move a lot when they’re awake during the night (e.g., extremely restless insomniacs, or frequent bathroom users) in your participant pool.1
What this easy-to-classify wake does is juice the numbers with a lot of easy dunks. If I only have three minutes of true wake in my night, and I misclassify one of them as sleep, I’m down to 67% specificity, all from that one wrong minute. If I have 30 minutes of true wake and 27 of them are super easy-to-classify—unambiguous, I’m walking around, no classifier on the planet would say this is sleep—then my specificity is going to be at least 90%, even if I get all three hard-to-classify minutes wrong.
So accuracy is no good and the usefulness of sensitivity and specificity hinges on the participant pool you use for testing. It gets worse. Most validation studies are, what, like thirty people? Assume you’ve got a big population, with a lot of spread in the amount of easy-to-classify wake epochs over the course of the night; something like this:
Assume also that you’ve got an algorithm that gets all the easy-to-classify wake epochs right, and all the hard-to-classify (motionless) epochs wrong. Now pick thirty people at random from your big population. How good does your classifier look? Here’s nine runs of this:
In one of your universes, the model looks amazing: 71% specificity! In another, it’s a terrible 22%. In all cases, it’s the same model. All that’s changing is who you tested it on. And this spread will only get bigger if you systematically select for healthy people in one study and people with a sleep disorder in another.
Of course you can imagine a generalization of this where instead of “sleep vs wake” you’re doing sleep staging—light (bundled N1/N2), deep (N3), REM sleep and wake.
I’m writing all of this because of the sleep algorithm lawsuit that dropped this week, in which Oura is accused of overstating its ability to match lab sleep tracking, or polysomnography (PSG)2. When I saw the headlines, my first thought was: “Wow, that’s weird. Oura’s really good.” (It is; I like them a lot).
As far as I can tell, the claims are:
Wristworn sleep tracking is not as good as EEG (✅, yes obviously; sleep stages were defined by EEG, the wrist is an imperfect proxy for the brain)
Wristworn sleep staging by Oura is basically a coin flip in terms of accuracy (❌, see below)
Oura did a bad thing by saying they had “95% Sleep Staging Accuracy” in marketing materials (🟡, yeah, they shouldn’t have done this— the defense is that they meant “two-class sleep staging” as in “wake vs sleep,” and not “light/deep/REM/wake sleep staging.” But, c’mon. I’m never going to assume two-class is what you mean when you say “Sleep Staging Accuracy.” And yet this also seems to me to be a clear case of Marketing Team did a Bad Ad, not some grand corporate conspiracy.)
Oura did a bad thing by reporting 79% agreement with PSG from a peer-reviewed study (❌, see below)
On the coin flip claim: What?? No! The coin flip claim comes from this paper where Oura was found to have 53.18% accuracy at four-class (light/deep/REM/wake) sleep staging in a clinical population (people with suspected sleep disorders, not healthy sleepers). But that’s not a coin flip when you’re trying to predict four classes. Consider if there existed 100 stages of sleep, and you picked them randomly over the course of the night. You would only expect to get 1% of the night correct! If you got 50% of the night right, it would be a sign that your algorithm is quite good at picking up on real signals.
Yes, okay, you could frame it as “for any given chunk of time, it was a coin flip as to whether or not it was the right or wrong stage.” But I don’t think people have good intuition for what 50% accuracy on a night of sleep looks like. It can look like this:
It’s not great, but it’s not terrible. You’re clearly picking up something in the signals.
Speaking of real signals: Have you ever seen a person score a night of sleep? I remember the first time I did. “Yeah, that looks mostly like REM,” the technician said, as we watched brain waves swim across the screen. “Eh, their eyes aren’t moving, but the sawtooth waves are there…” He shrugged and made the best judgment call he could for that 30 second chunk of time.
This is the gold standard! Of course you’re not going to be able to nail it!
For a quantification of just how squishy this all can be, see this paper, in which six human scorers tackling five-class sleep classification (N1/N2/REM/Deep/Wake) agree on only 46% of the night. It’s not that they’re bad at their job! It’s that we’re trying to fit a complex, high-dimensional process that’s nonuniform across the brain into a bucket of labels that you can count on one hand. If someone’s partially in REM, partially in deep sleep, they have no way to express that. It ends up being a true coin flip. And a sleep tracker’s accuracy can be punished for coming down on the other side.
Finally, on the 53.18% vs 79% accuracy: Yes, 53% is low. But that was in a population with sleep disorders, not healthy people. And even if it wasn't a different patient population, it’s important to remember the lesson from eight paragraphs up. Just as two random groups of participants can give you a markedly different picture of sleep vs. wake accuracy, so too could two random groups of participants give you a markedly different picture of sleep staging accuracy3. It doesn’t mean you did something wrong. You cast your fishing line into the pond four times, and in three studies, you got a catfish, and in the other, you got a minnow.
Where the analogy breaks down is that fishing is cheap, so you could cast your net out over and over again in an effort to arrive at a conclusive quantitative treatment of “the probability of getting a catfish from this pond.” Sleep studies are expensive, so you’ve basically got just a handful of lines to throw out there. As a non-lawyer, I can’t really say how much Oura should have added “*in healthy people” caveats to their advertising. But I don’t see an issue with them using 79% in their ads. The catfish still happened. You still caught those fish.
P.S. I’ve got my own company, Arcascope, but I have no relationship with Oura and nobody asked me to write this.
No one will catch you if you do this and never share the data.
Oura’s response is here: https://ouraring.com/blog/how-oura-measures-sleep-and-validates-accuracy/
If this were a fight about algorithms spawning from academia, I’d say guys, just use a benchmark dataset. When the participant pool can change your numbers this much, you should remove participant pool as a variable. You should run your algorithm on open source datasets (like mine!) of raw signals and co-recorded PSG, and update your performance numbers on that same dataset as your algorithm improves. When you use a benchmark, it doesn’t matter how many people had lots of easy-to-classify wake or how many people had ambiguous sleep stages throughout the night: you’re reporting numbers for that set of people, which can be fairly compared to other numbers for that set of people.
The rub is that consumer wearables each have their own sensors, so you can’t just use (say) a benchmark dataset of Apple Watch heart rate to report metrics for Oura’s sleep algorithm. That said, Oura could release an open source dataset of acceleration and heart rate raw data from their devices, along with co-recorded PSG, and let researchers at it. People would squeeze so many papers out of it for free.





