EEG, eye movement, and muscle tone. A clinical night is a different kind of fact.
Setting the next column
CCSetting the next column
CCCulture Analysis / Strong
A sleep lab records brain waves, eye movements, and muscle tone overnight. A wrist sensor records pulse and motion, then an app compresses that into a number that looks like a lab result. The number can be stable from night to night without being the same kind of fact.

The argumentScores feel like medical facts because they are delivered with the visual authority of a clinical readout, on the same devices that sometimes do carry FDA-cleared functions such as an electrocardiogram. What moved from the lab to the wrist is not polysomnography. It is a cheaper sensor stack (optical pulse plus motion) plus a proprietary index, released into a regulatory gap the FDA named in 2016 and 2019: low-risk 'general wellness' products are not examined as medical devices if they do not claim to diagnose. A score can be internally consistent, the same algorithm, the same person, night after night, without being a clinical measurement of sleep architecture or recovery.
Why do recovery scores and sleep stages feel like facts, and what actually changed when tracking moved from labs to wrists?
A 2022 sleep-lab study of six consumer devices found 86 to 89 percent agreement with polysomnography for sleep versus wake, and only 50 to 65 percent agreement when the task was naming a specific sleep stage. Detecting that you were down for the night is a different job from staging the night.
The FDA's general-wellness policy, restated in 2016 and 2019 and updated in 2026, leaves low-risk products that promote a healthy lifestyle outside ordinary medical-device review, as long as they do not claim to diagnose, treat, or substitute for a cleared device. Sleep stages and recovery scores typically live in that bucket.
The same Apple Watch that received De Novo clearance for an ECG app in 2018 can also display a sleep score that is not that kind of clearance. One wrist, two regulatory categories, which is part of why the score inherits the ECG's aura.
Kelly Baron and colleagues coined 'orthosomnia' in 2017 for patients who trusted tracker data over how they felt, and sometimes over a sleep study. The term is a case-series label, not a DSM diagnosis. It names a cultural problem the devices created: a number that is easy to obsess over without being easy to validate.
A clinical sleep study, polysomnography, is an overnight recording of brain waves, eye movements, muscle tone, breathing, and often blood oxygen, scored by a technician using agreed rules. It is expensive, inconvenient, and still the reference standard against which consumer gadgets are judged. A recovery score is a different object. A ring or a strap measures pulse with light and movement with an accelerometer, then an unpublished formula emits a number, often out of 100, that is easy to read before coffee.
The number feels like a vital sign because it is presented like one: a single index, a color, a trend arrow, sometimes a comparison to 'your baseline.' What changed is not that sleep science discovered a new stage of rest. What changed is that a lab procedure was replaced, in daily life, by a wrist sensor and a score, under a federal policy that treats most of those scores as wellness accessories rather than medical devices.
EEG, eye movement, and muscle tone. A clinical night is a different kind of fact.
Pulse and motion, compressed into a score with the visual authority of a readout. Stable is not the same as equivalent.
The cultural name for the earlier, nerdier version of this habit is dated. Gary Wolf and Kevin Kelly founded Quantified Self in 2007, a meetup culture of people who logged weight, mood, location, and sleep in spreadsheets and early gadgets, aiming at 'self-knowledge through numbers.' That project was voluntary, small, and often self-skeptical. The home version that followed is default software on a watch. You do not join a meetup. You wake up to a score you did not ask a clinician to interpret.
This is a different argument from the one about the sleep-product market. Mattresses, melatonin, and smart beds are a retail category. The quantified-self problem is epistemological: why a proprietary index is treated as a measurement of the body, and what the lab-to-wrist swap actually discarded.
| Element | Overnight polysomnography | Typical consumer wearable |
|---|---|---|
| Brain activity | EEG electrodes. Sleep stages are defined from brain-wave patterns, not from pulse. | None. Stages are inferred from motion and optical pulse, then labeled with the same words (light, deep, REM). |
| Eyes and muscle tone | EOG and EMG, which help separate REM from other stages. | Not measured. |
| Heart | Often ECG leads, a direct electrical signal. | Photoplethysmography (PPG): light shone into skin, pulse inferred from blood-volume change. Fine for many trends, not the same signal as ECG. |
| Output | A scored hypnogram plus respiratory events, produced under clinical protocols. | Minutes per stage, a sleep score, sometimes a recovery or readiness index with unpublished weights. |
| Regulatory status | A medical procedure in a regulated lab. | Usually a general-wellness feature, unless the company seeks clearance for a specific diagnostic claim (for example an ECG app). |
Validation studies keep finding the same split. In 2022, Miller, Sargent, and Roach put 53 healthy adults in a sleep laboratory for one night with six consumer devices and with polysomnography. For the two-state question, asleep or awake, agreement with the lab ran from 86 to 89 percent. For the multi-state question, which stage, or wake, agreement fell to 50 to 65 percent, with Cohen's kappa values that the authors read as slight to moderate. Their conclusion was specific: the devices were reasonable for field estimates of when sleep happened and how long it lasted, and not yet adequate for staging.
A 2024 Brigham and Women's study in Sensors compared Oura Ring Gen3, Fitbit Sense 2, and Apple Watch Series 8 against polysomnography in 35 adults without a diagnosed sleep disorder. Sleep-versus-wake sensitivity was 95 percent or higher for all three. Stage sensitivity ranged from 50 to 86 percent depending on device and stage. Oura was not statistically different from the lab on wake, light, deep, or REM duration in that sample. Fitbit and Apple Watch were, in opposite directions on light and deep sleep. The paper discloses funding from Oura Ring Inc. That does not make the measurements fake. It does mean the head-to-head should not be read as an independent ranking of brands, which this article will not provide.
The devices are much better at noticing that a night of sleep occurred than at naming the stages a recovery or sleep score is often built from. Consistency across your own nights is not the same as agreement with a lab.
Miller, Sargent, and Roach, Sensors 2022, DOI 10.3390/s22166317 · accessed 2026-08-12 · 53 healthy adults (mean age 25.4), one laboratory night, 9 hours in bed. Each participant wore six devices. Reference standard: polysomnography for sleep, ECG for heart rate. Agreement is epoch-level percent agreement with PSG, not a consumer 'accuracy' marketing claim. Devices are 2022-era hardware (Apple Watch Series 6, Oura Gen2, WHOOP 3.0), so later algorithms may differ. The pattern, two-state far above multi-state, has been replicated in later studies.
Recovery and readiness scores add a second compression. They typically mix overnight heart-rate variability, resting heart rate, and sleep duration or staging, then weight those inputs in a way the company does not publish as a clinical formula. Heart-rate variability from a good optical sensor can track in the same direction as ECG-derived HRV for many people, which is useful for noticing a trend after a hard week. It is not a diagnosis of overtraining, illness, or 'recovered enough to race.' WHOOP built a business on that index as a membership. Oura and Apple display adjacent scores. The cultural move is the same in each case: several noisy signals become one authoritative integer.
An integer can be reliable in the statistical sense of repeating itself. Wear the same ring, sleep in the same room, get 82 three days running. That reliability is what makes the number feel scientific. Clinical meaning is a different test: does 82 correspond to a construct a sleep physician would recognize, in a population that includes insomnia, apnea, and shift work, not only healthy volunteers in a lab? The validation papers above were run in healthy adults. They do not license the score as a fact about disease.
The regulatory gap is not a loophole someone forgot to close. It is written policy. In 2016, section 3060 of the 21st Century Cures Act removed certain software functions intended only to maintain or encourage a healthy lifestyle, unrelated to diagnosis or treatment, from the definition of a device. The same year, FDA's device center said it did not intend to examine low-risk general wellness products for device compliance. The 2019 guidance restated the two-part test: general wellness intended use, and low risk (not invasive, not implanted, no inherently dangerous energy). In 2026 the agency updated the policy again to address wearables that estimate physiologic parameters such as blood pressure, oxygen saturation, glucose, and heart-rate variability. Those products can still be wellness products if they stay on the wellness side of the line: not for diagnosis, not a substitute for a cleared device, not for medical management.
That 2026 update matters because it is the opposite of a crackdown. As sensors got closer to clinical quantities, FDA clarified that some of those quantities can still be sold as lifestyle numbers. The score culture did not outrun the regulator by accident. It grew inside a category the regulator defined.
Gary Wolf and Kevin Kelly start a community around self-tracking tools. The ethic is n-of-1 experiment, not a mass-market morning number.
Congress carves certain healthy-lifestyle software out of the device definition. FDA states it will not examine low-risk general wellness products as devices.
Baron and colleagues describe patients who present with tracker-driven sleep anxiety, sometimes preferring the device to a sleep study. The paper estimates that about 10 percent of US adults then used a wearable tracker regularly.
DEN180044 puts a medical-grade function on a consumer watch, tightening the association between the wrist and clinical truth.
FDA's update explains how non-invasive estimates of blood pressure, oxygen, glucose, and HRV can remain general wellness if they do not cross into diagnosis or substitute for cleared devices.
Quantified Self; Federal Register 2016-17902; Baron et al., J Clin Sleep Med 2017; FDA DEN180044; CRS IN12657 on the 2026 guidance.
Population sleep, measured the old way, did not become a solved problem in the years the scores proliferated. The CDC's National Health Interview Survey found that in 2024, 30.5 percent of US adults averaged less than seven hours of sleep in a 24-hour period, the American Academy of Sleep Medicine's recommended minimum. Only 54.8 percent said they woke feeling well-rested most days or every day. Those are survey answers, not ring data. They are also a reminder that a culture of nightly scores can expand while the share of people who sleep enough does not move in any dramatic way.
In 2017, Kelly Baron's group at Rush and Northwestern described patients who arrived at clinic because a tracker said their sleep was inadequate, including people whose lab studies did not match the device's story. They called the pattern orthosomnia, on the model of orthorexia: a perfectionist quest for 'correct' sleep. It is not an official diagnosis. It is a clinical observation that the score can become the symptom. The quantified self at home is that observation at population scale. The lab never sent a 0-100 recovery index home. The market did, under a wellness policy that asked the score not to claim it was medicine, and did not ask it to stop looking like medicine.
Treat night-to-night direction (much worse than your usual) as a prompt to notice alcohol, illness, or a schedule change. Do not treat a stage minute-count as architecture a sleep lab would sign.
A cleared function (ECG, a specific apnea feature) and a wellness score can share a device. Check which claims the company is actually allowed to make.
If the number is making sleep worse, that is a documented pattern, not a personal failure of discipline. Baron's 2017 cases were people who did what the interface invited: they believed it.
Consistency means the algorithm is stable on you. Clinical meaning means the algorithm matches a construct measured another way, in the people who have the problem. You can have the first without the second. A bathroom scale that is always two pounds heavy is consistent. It is still wrong about weight.
Not on the evidence here. Duration and timing are the parts that track the lab most closely, and they are the parts sleep-regularity research actually cares about. The part to downgrade is the staged, scored, 'recovered' narrative.
No. The wellness policy is about not running them through medical-device review when they stay on the lifestyle side of the line. Unsafe would be a different finding. The cultural risk is misplaced certainty, not a shocking from the charger.
Sources and further reading
53 adults, one night in a sleep lab, six devices versus PSG and ECG. Two-state sleep/wake agreement 86-89 percent (kappa 0.30-0.51). Multi-state stage agreement 50-65 percent (kappa 0.20-0.52). Authors conclude devices are valid for field timing/duration of sleep but require improvement for specific stages.
35 adults, Oura Ring Gen3, Fitbit Sense 2, Apple Watch Series 8 versus PSG. Sleep-versus-wake sensitivity at or above 95 percent for all three. Stage sensitivity 50 to 86 percent depending on device and stage. Study received funding from Oura Ring Inc. Cited with that limitation.
CDRH does not intend to examine low-risk general wellness products to determine whether they are devices, or whether they comply with device requirements, if they are intended only for general wellness use and present a low risk (not invasive, not implanted, no inherently risky intervention).
Summarizes the 21st Century Cures Act (2016) software carve-out, the 2019 wellness guidance, and the 2026 update addressing wearables that estimate physiologic parameters (blood pressure, oxygen saturation, glucose, HRV) when intended solely for wellness and not as substitutes for cleared devices.
De Novo classification of the Apple ECG app as a medical device, establishing that some watch functions are cleared as devices while adjacent wellness scores on the same hardware are not.
Case series coining orthosomnia: preoccupation with perfecting wearable sleep data. Notes then-current estimates that about 10 percent of US adults used a wearable tracker regularly and 50 percent would consider buying one. Not a DSM diagnosis.
2024 NHIS: 30.5 percent of adults averaged less than 7 hours of sleep; 54.8 percent woke well-rested most days or every day. Population sleep, measured by survey, not by wearables.
Wolf and Kevin Kelly founded Quantified Self in 2007 as a collaboration of users and makers of self-tracking tools seeking 'self-knowledge through numbers.'

Continue reading
CISA's four habits (multifactor, updates, think before you click, a password manager) only hold if someone in the house can recover the others. Shared streaming logins, one spouse's email as the family reset, and a router nobody else can log into are how a personal checklist fails as infrastructure.
Culture / Digital Life / 8 min
READ NEXTAbout the byline
Sources, review method, and commercial relationships appear with each story. Reach the desk at editorial@culturecolumn.com.