SSMT Claim Register
SSMT-2026-0002v12026-07-26 Type M · Measurement Wearables & Recovery draft
Claim, as marketed
“See your deep, REM and light sleep every night.”
Canonical claim
Consumer wearables measure sleep stages (light/deep/REM) accurately.
Scope
Consumer wrist and ring wearables. Applies to stage classification; sleep-wake detection and total sleep time are better supported.
Verdict
Partially supported Directionally useful in aggregate, unreliable for any single night — and the accuracy that is well established is a different measurement from the one being sold.
Real evidence, narrower than the claim.
Claim drift — 5 of 7
PopulationSettingComparatorOutcomeMagnitudeReproducibilityIndependence
Outcome drifts because the largest, strongest evidence (n=798) covers total sleep time and explicitly did not analyse staging at all — yet staging borrows its credibility.
Evidence located
10
studies located
8
independent
9/10
used PSG criterion
62
largest independent primary n
Findings

Sleep-wake detection is reasonably good, with one systematic flaw. Sensitivity is consistently 0.91–0.99 and overall accuracy 0.85–0.90. But specificity — correctly calling wake 'wake' — is 0.18–0.54. Devices are biased toward calling everything sleep. Pooled total sleep time error across 24 studies (n=798) was −16.9 min.

Stage classification is substantially worse. Four-state Cohen's κ sits at 0.21–0.53 in healthy and mixed adults, and 0.06–0.31 with 35–53% raw epoch accuracy in clinical sleep-lab patients. Per-stage sensitivity for the two stages marketing sells hardest: deep 0.45–0.75, REM 0.33–0.87.

The trap. Night-total minute comparisons look excellent — one meta-analysis found light −4.3, deep +1.4 and REM −3.9 minutes, none significant — while epoch-level accuracy sits near 53%. Over- and under-calls cancel out in the nightly total. A device can report '90 minutes of deep sleep' correctly on average while placing it in the wrong 90 minutes.

Why reproducibility drifts. Group-level agreement masks large individual-level error, and the bias is proportional — so it cannot be calibrated away. The documented failure mode across three independent groups: wake, deep and REM all collapse into 'light', which is the algorithms' default under uncertainty.

Worth watching on independence. Scored aligned because strong independent PSG validations exist. But the most favourable staging numbers come from vendor-funded and vendor-authored work (specificity 73–75%) while independent studies report 29–54%.

What would change this verdict
An independent, multi-night, at-home ambulatory-PSG validation, n≥150, in a demographically and clinically diverse sample, reporting epoch-by-epoch four-state κ ≥0.6 with per-stage sensitivity ≥0.80 for both deep and REM, and reporting individual-level limits of agreement rather than group mean bias.
Practitioner read
Your watch is good at telling when you are asleep versus awake, and total sleep time is usually within about 15–20 minutes of a lab study — trust that trend. The deep and REM breakdown is a rough estimate, not a measurement: against clinical sleep studies these devices put sleep in the right stage roughly half to two-thirds of the time, and they are worst at noticing you are actually awake in bed. Use week-to-week changes in your own numbers; read nothing into a single night's deep-sleep figure.
Could not be verified
Listed rather than dropped. Counts are floors, not censuses.
Sources
  1. Chinoy ED, et al. (2021). SLEEP 44(5):zsaa291
  2. Schyvens A-M, et al. (2025). SLEEP Advances 6(2):zpaf021 — 6 devices, n=62
  3. Herberger S, et al. (2025). Sci Rep — clinical sample, n=45
  4. Lee YJ, et al. (2025). J Clin Sleep Med — meta-analysis, n=798
  5. Haghayegh S, et al. (2019). J Med Internet Res 21(11):e16273
Every source must resolve at a DOI or PubMed ID.
Changelog
Cite this card
SSMT. "Consumer wearables measure sleep stages (light/deep/REM) accurately." Claim SSMT-2026-0002 v1, graded 2026-07-26. https://sportssciencetech.com/c/SSMT-2026-0002
Disagree with this grade? Every card carries “what would change this verdict” so the specification is explicit. Send evidence and it will be logged and the card re-versioned in public — including when it moves a grade in a manufacturer's favour. Submit evidence →