I tore my left ACL skiing in February 2023 and did not find out for six months. My phone knew something was wrong the next morning. Then it stopped knowing. So I went back through two years of passive wearable data, with my own clinical record as the referee, to see what it saw and what it missed.
The setup
The fall itself was ordinary. I went backward with the skis still attached and the knee tangled behind me. Ski patrol brought me down. I used a cane for a week and could not fully bear weight for about two. After that, life felt normal, except that I could not push off or hop on the leg. Nobody examined the knee until August, when I asked for imaging because I wanted to ski that winter. An X-ray and then an MRI found a chronic complete ACL tear, a displaced bucket-handle tear of the lateral meniscus, and a cartilage defect on the tibial plateau. Two hands-on exams that same fortnight had come back negative for the ACL. Surgery was in September 2023. The first run back was 349 days later. The first ski day was 812.
Through all of it I was wearing an Apple Watch and carrying an iPhone. An Oura ring arrived one month before surgery. The Slopes app logged every ski run with GPS. From December 2024 the Strong app logged every gym set, left leg and right leg separately. None of that was collected for this project. It was just there.
What I was actually asking
The obvious question is whether wearables can detect an injury. I did not think that was the honest question. I already knew the answer was "sort of." The better questions were narrower. Could rules written only against healthy data have flagged the tear, and for how long? Which was the bigger physiological event, the tear or the surgery? Did recovery follow an order, and did the metrics agree with the rehab protocol? And where the data and my own memory disagreed, which one was right?
That last one shaped the whole project. I remembered walking normally within weeks. I remembered sleep broken by knee pain that winter. Both of those turned out to be partly wrong, in different directions. I only know that because the data was sitting there the whole time.
Getting the data to agree on what a day was
Most of the work went into making six sources describe the same days.
The iPhone gait metrics (walking speed, step length, double support, asymmetry) came from the native Apple Health export. The Watch supplied steps, flights, stair speeds, resting heart rate, HRV and sleep. Oura supplied readiness, sleep architecture and skin temperature. The first problem was that Oura also writes into Apple Health. Once the ring arrived, Apple's sleep and heart rate tables had two authors. I had already been burned by this in an earlier comparison study. So every Oura-written row was stripped out of the Apple data before anything else happened. Otherwise the ring's arrival, one month pre-op, would have looked like a physiological change timed suspiciously close to surgery.
The second problem was hardware. The Watch was a Series 6 through 2022 and an Ultra from January 2023. That means stair speeds, flights and step counts are not like-for-like across the two years. I flagged this everywhere it mattered. In one case it turned out to explain a "deficit" outright. The phone was the same 13 Pro Max throughout. Gait is the clean series.
Everything was collapsed to one row per day, then to weekly medians for the charts. Days were then labelled by phase using the clinical timeline rather than the calendar. Pre-injury, undiagnosed, pre-op, post-op, PT block, comeback.
The clinical record as referee
The part I am most glad I did was pulling the clinical record in as data rather than as background. Apple Health's native export includes the FHIR records from the hospital. Surgery dates, imaging orders and medication orders came out with timestamps. I added the visit notes from the patient portal by hand. That gave me 37 documents and a dated list of every PT visit, orthopaedic follow-up, clearance, protocol transition and one setback.
This changed the analysis from "look for interesting shapes in the data" to "here is a list of events with dates, does the data see them?" Some of those events were things I had forgotten. The setback in March 2024 is in a physical therapy note, not in my memory. The data found it to the day.
Testing the tear like it was still unknown
To answer whether the tear was detectable, I could not just look at the post-injury data and point at the drop. Of course it drops. I already knew where to look.
So I wrote alert rules using only pre-injury data as their reference. One rule fires when a weekly median sits beyond the pre-injury 5th or 95th percentile for three consecutive weeks. A composite rule fires when three or more gait metrics are abnormal for two weeks running. Then I ran those rules on the injured window, February to August 2023, and on the same calendar dates in healthy 2022 as a control.
The control is what makes it honest. The composite rule fired on day 15 in 2023 and never in 2022. Six of eight gait metrics broke the band for three or more weeks in 2023. None did in 2022. Then every alert switched off within about two months. From weeks 9 to 25 the deficits were still there statistically (Cliff's delta 0.25 to 0.40) but inside the normal band. A threshold system would have said "injured, then recovered." So did two clinicians' hands. Only the MRI said what it was.
Comparing the tear to the surgery
Before running anything, I wrote down a prediction. The tear would hit gait harder, the surgery would hit physiology harder. Writing it down first meant I could not quietly revise it afterward.
The comparison was days 0 to 6 after each event against a 31-day baseline before it. I standardised everything as Cohen's d so the two events could sit on one axis. Only Watch and iPhone metrics were comparable. There was no Oura in February 2023. The prediction failed. Surgery was the larger perturbation on every comparable metric, including gait. The vivid physiological response to surgery, readiness down 3.4 standard deviations and sleep efficiency from 88 to 61 percent, was Oura-only. The Watch barely registered it. That half of the prediction could not be tested symmetrically. That is worth saying out loud rather than folding into a tidier story.
Time to recovery, and a correction I had to make
For each metric I measured the number of days after surgery until the weekly median re-entered the pre-injury interquartile band. This gave recovery an order. Double support first at 59 days, walking speed last at 199.
Walking speed had a trough from January to March 2024 that I first attributed to winter. That felt obvious. I tested it anyway by fitting a seasonal baseline (a Fourier fit on three prior winters, 2020 to 2023) and subtracting it. The pre-injury seasonal cycle is real but small, about 5 percent. It explained 6 percent of the 2024 trough. The rest lined up with the months physical therapy lapsed. My first reading was wrong. The correction also cut the other way for the 2023 undiagnosed months. Those were spring and summer, so the persistent deficit there got larger after adjustment, not smaller.
Finding things nobody told the data about
The setback was the best test of the whole approach. A March 2024 clinic note describes a pivot, a pop, and pain going downstairs "about two weeks ago." I took the 14 days before and after and compared every metric. Stair descent speed fell from 1.60 to 1.26 feet per second, starting on March 10. Nothing else in gait moved. Resting heart rate improved. Oura saw it too, as awake time doubling for two weeks while readiness rose. That is a local pain signal, not a systemic one. It is also the only sleep fragmentation event in the record after the first two post-op weeks.
The other direction was just as useful. A December note says I was waking from knee pain. Oura's October to December sleep was as consolidated as the recovered year. What was abnormal was sleeping heart rate and HRV, every month until June. That is deconditioning, or recovery. The data cannot separate the two. But it is not "broken by pain."
What passive data could not see
I kept a running list of things the data missed. A project like this is only useful if the negatives are as visible as the hits.
PT starting on December 12 has no counterpart in any signal. Strength gains, measured in the clinic and later in the gym, never showed on the stairs. Hop and range of motion, the two things I actually felt, have no sensor proxy at all. The gym log from Strong gave me a limb symmetry index across 40 paired sessions. No pre-injury baseline exists, though. The right leg is the comparator, and it was detraining at the same time. I had phone sensor data from Beiwe, the research platform I work on. It begins six months after surgery and was not used.
And the statistics. This is a single subject. Daily values are autocorrelated. That makes every p-value optimistic. The numbers to quote are effect sizes with bootstrap confidence intervals. That is what the results page shows.
Where the numbers are
Every chart, table and comparison lives on the results page. It runs chronologically through the injury, the surgery and the recovery, with separate views for running and skiing. The last of those is the one I find hardest to look at. The comeback season was the same runs, with the same vertical per run, skied slower and half as many times a day. Top speed went from 33 to 27 miles an hour. It never ramped up within the season. Morning-after asymmetry stayed at 0 to 5 percent every time. The knee was fine. I was not.