Your Apple Watch already knows which lift you're doing. This teaches it to write it down.
Automatic exercise recognition and rep counting from wrist IMU, running entirely on-device — with a companion iPhone app where every correction you make becomes tomorrow's training data.
How it works · The models · Data · The UI · Install · Training
Both demos are animated SVG mockups built from the real view code and design tokens (docs/make_demo_svgs.py), not screen recordings — the layout, copy and palette are lifted from LiftLoggerWatchApp.swift, SessionsView.swift, SessionDetailView.swift and DesignSystem.swift. To swap in a real capture, drop a .gif in docs/ and change the <img src> above.
- What it does
- How it works
- The models
- Data & how it is loaded
- The UI
- The correction loop
- Installation
- Training recipes
- Repo layout
- Limitations
- How this was built
You start a session on the watch and lift. That's the whole interaction.
- No exercise selection. A 1-D CNN reads the wrist IMU and names the movement ~2× per second.
- No rep counting. When a set closes, one of three rep models counts it from the same signal.
- No manual sync.
WatchConnectivityships the session's CSVs to the phone at End Session. - No wasted correction. Fixing a wrong count on the wrist or the phone marks that set
reps_confirmed, and confirmed rows are exactly what the next training run reads. - Nothing leaves the devices. Core ML on the Neural Engine; no server, no account, no network calls.
On the watch — from wrist motion to a logged set:
flowchart TD
A["CMMotionManager · 50 Hz<br/>acc_xyz + gyro_xyz"] --> B["rolling 2 s window<br/>100 samples, re-run every 25"]
B --> C{"CNN1D classifier<br/>~2 predictions / sec"}
C -->|"confidence below 0.6,<br/>or the class is still unstable"| B
C -->|"same class 3 windows running"| D["SET OPENS<br/>buffer the bout + 3 s of lead-in"]
D --> E{"rest stable 5 windows<br/>and 5 s grace elapsed?"}
E -->|no| D
E -->|yes| F["SET CLOSES · hand the bout to a rep counter"]
F --> G["RepCounter<br/>autocorrelation"]
F --> H["RepDensityCNN<br/>density curve"]
F --> I["RepPeriodCNN<br/>rep period"]
G & H & I --> J["detected.csv · reps_confirmed = 0"]
Off the watch — the flywheel that makes the next model better:
flowchart LR
A["detected.csv<br/>reps_confirmed = 0"] -->|WatchConnectivity| B["iPhone · Sessions"]
B --> C["you fix or confirm<br/>−/+/✓ on the wrist,<br/>Fix Reps on the phone"]
C --> D["reps_confirmed = 1"]
D --> E["Build merged export<br/>readings · sets · reps"]
E --> F["training/<br/>PyTorch"]
F --> G[".mlpackage"]
G -->|"drop into the Watch target"| H["a better model<br/>next session"]
The loop closes: the sets you correct are the sets the next model learns from.
Every constant that has to agree across the Swift and Python sides is named in both files with a
MUST match comment — WINDOW/windowSize, FS/fs, REP_WIN_STRIDE/repWinStride,
REST_LABEL/restLabel, DISCARD_LABEL/discardLabel.
Every knob below lives at the top of RecorderModel.swift with the reasoning attached.
| knob | value | why |
|---|---|---|
| sample rate | 50 Hz | matches config.FS; fast enough for a 0.5 s rep, cheap enough to run all session |
| inference window | 100 samples (2 s) | matches config.WINDOW — a couple of reps of context |
| inference stride | 25 samples | a prediction ~2× per second |
| confidence floor | 0.60 | below this the window counts as rest; don't log guesses |
| windows to open a set | 3 | debounce, so one lucky window can't start a set |
| windows to close a set | 5 | deliberately harder to leave an exercise than to enter one |
| rest grace | 5 s | a hold at the top, a breath, a grip reset all look like rest to a 2 s window — this keeps the set from fragmenting |
| minimum set | 1.5 s | anything shorter is noise, not a set |
| bout lead-in | 150 samples (3 s) | kept before the set commits so the first reps aren't clipped |
| bout cap | 6000 samples (120 s) | hard memory ceiling per set |
| countdown | 3 s | manual sessions only — time to get into position before start_ms is stamped |
Four models total: one that answers what you're doing, and three competing answers to how many. All of them are the same shape of thing — small 1-D convnets over 6-channel IMU — and all of them convert cleanly to Core ML and run on the Watch's Neural Engine.
training/models.py · trained by train.py · exported by export_coreml.py
input (1, 100, 6) raw acc_x/y/z, gyro_x/y/z @ 50 Hz — the model normalizes internally
conv7 → 64 → BN → ReLU → maxpool2
conv5 → 128 → BN → ReLU → maxpool2
conv3 → 128 → BN → ReLU → global average pool
dropout 0.3 → linear
output softmax over N exercises + "rest" ~150k parameters
Why a temporal ConvNet and not something else:
- Translation invariance. A 2 s window can start anywhere in a rep — at the top, mid-descent, in the pause. Convolution + global average pooling means the model doesn't care.
- Small-data robustness. Global average pooling instead of a flattened dense layer keeps the parameter count at ~150k, which matters when a class has 40 examples, not 40,000.
- It deploys. No recurrence, no attention, no per-prediction search over a training set — a
clean
torch.jit.trace→coremltoolsconversion that the Neural Engine likes.
The normalization statistics are computed on training windows only and baked into the exported
graph (ExportWrapper), so the Swift side hands over a raw sensor buffer and reads back
probabilities. There is no feature engineering to keep in sync across two languages.
A Random Forest on hand-crafted window features (baseline_rf.py, features.py) is kept around as
the number to beat. If the CNN can't clearly beat it, the bottleneck is data, not architecture.
training/README.md has the full ranking of alternatives considered (DeepConvLSTM, TCN, k-NN+DTW,
transformers) and why each was accepted or rejected for a watch.
The 15-exercise roster (RecorderModel.exercises, mirrored by config.KEEP_EXERCISES):
incline chest press · machine chest press · machine shoulder press · wide-grip machine row ·
cable push down · overhead triceps · dumbbell hammer curl · dumbbell curl · cable curl ·
forearm raises · lat pulldown · squat · dumbbell RDL · machine calf raise ·
dumbbell Bulgarian split squat — plus an implicit rest class for everything else. The phone's
label picker (SessionStore.knownExercises) additionally lists a few retired names, so old
sessions recorded before the roster changed can still be relabeled.
Note on the bundled checkpoint. The
.mlpackagecommitted inLiftLogger Watch App/was exported when the roster was 7 exercises +rest(cable_push_down,dumbbell_hammer_curl,forearm_raises,incline_chest_press,machine_row_wide,machine_shoulder_press,overhead_triceps). The watch UI andconfig.KEEP_EXERCISESnow offer 15 — retrain and re-export to cover the rest.
training/reps.py ⟷ RepCounter in RecorderModel.swift · evaluated by evaluate_reps.py
A rep is one cycle of a quasi-periodic movement, so no learning is required at all:
- For each of the 6 channels, compute the normalized autocorrelation over lags in
0.5–4.0 s (
REP_PERIOD_RANGE_S, i.e. 0.25–2 reps/sec). - Keep the channel with the strongest peak — the axis that best expresses this exercise picks itself, which is why one implementation covers curls and squats alike.
- Its peak lag is the rep period
T;reps ≈ bout_length / T. - Cross-check with a peak count on that channel (minimum spacing
0.6·T, prominence0.30 × signal std). If autocorrelation is weak (< 0.25), trust the peaks instead; if the two agree within 1, prefer the peaks — they handle partial first/last reps slightly better.
| Needs to run | nothing — no weights, no bundle, always available |
| Needs to improve | nothing trainable; evaluate_reps.py fits a per-exercise linear correction true ≈ a·pred + b from confirmed counts (≥ 5 sets per exercise) and writes artifacts/rep_calibration.json |
| Strength | zero cold-start, fully interpretable, identical logic in Python and Swift |
| Weakness | a systematic bias per exercise, and it degrades where the wrist barely moves periodically |
This is the floor every learned counter has to beat, and the fallback at the end of every chain.
training/models.py · data by rep_windows.py · trained by train_reps_windows.py ·
exported by export_reps_coreml.py ⟷ RepDensityCounter in RecorderModel.swift
Instead of regressing a count, this predicts where the reps are and integrates:
input (1, 400, 6) an 8 s window at real 50 Hz
conv7 → 64 stem, then 7 dilated conv3 blocks, dilations (1,2,4,8,16,32,64)
NO pooling over time — output length stays 400
conv1x1 → Softplus
output (1, 400) per-frame rep density ≥ 0; reps = Σ density
The labels are per-rep tap timestamps from the phone's Rep Tagger (RepTapView): an observer
taps once per rep while you lift, each tap becomes a unit-area Gaussian (σ = 0.20 s) on the density
curve, so the curve integrates back to the true count. On the watch, RepDensityCounter slides the
same 8 s window across the bout at a 1 s stride and overlap-adds by averaging — dividing by how
many windows covered each frame, so a rep in an overlap region isn't counted twice — exactly
mirroring rep_windows.bout_count in Python.
Three design details that are easy to get wrong, all handled explicitly:
- Receptive field must span a rep.
RF = 7 + 2·Σdilations. The legacy whole-bout dilations(1,2,4,8)give 37 frames — fine when a frame is 1/256th of a set, but only 0.74 s in real time, less than one rep of a slow lift.REP_WIN_DILATIONSreaches 261 frames ≈ 5.2 s. - Head bias initialization. A real density averages ≈ 0.01 reps/frame, while an untuned Softplus
head starts at 0.69 — 70× too high. The trainer initializes the output bias to
inverse_softplus(mean target). Without it, the model burns its entire epoch budget just deflating, which looks exactly like convergence. - Windows, not whole bouts. An earlier path (
train_reps_model.py, kept for comparison) resampled each set to 256 frames — which made a 10 s set and a 50 s set at the same tempo look 5× different, and gave ~50 training examples where the classifier had thousands. Fixed-second windows turn one 30 s set into ~23 examples and keep tempo in real Hz.
| Needs to run | LiftLoggerRepCounter.mlpackage bundled (it is — the windowed variant) |
| Needs to improve | per-rep tap timestamps in reps.csv — dense, accurate, effortful: someone has to tap along |
| Strength | most accurate when it has data, and the only counter that tells you where each rep was |
| Weakness | the tapping bottleneck; every new exercise needs a live tagger session |
training/models.py · trained by train_reps_period.py · exported by
export_reps_coreml.py --period ⟷ RepPeriodCounter.swift
The counter built specifically to remove the tapping bottleneck. It trains on nothing but the
final rep integer already sitting in sets.csv:
input (1, 600, 6) 12 s of the bout, edge-padded if short / cropped from the start if long
same conv/pool trunk as CNN1D (global average pooled)
dropout → linear → 1 scalar
output log(period_seconds) → exported wrapper returns exp(·), i.e. SECONDS
the watch computes reps = round(bout_duration / period)
Two choices carry this model:
- Predict the period, not the count. A count entangles tempo with however long the set happened to run; the period is duration-independent and lives in a narrow physical band (0.5–4.0 s). That makes a single scalar per set enough supervision to actually learn from — a far more sample-efficient regression target.
- Train in log-space. Periods are ratio-scale: a 2 s rep vs a 1 s rep is "twice as slow", not "1 s slower". Log targets stop fast and slow reps from contributing wildly different-magnitude losses.
Its training set grows every time you use the app: every hand-dialed manual set, plus every
auto-detected set you nudged or checked off — on the wrist or in the phone's Fix Reps sheet.
Sets with fewer than 2 reps are ignored (REP_PERIOD_MIN_REPS).
| Needs to run | LiftLoggerRepPeriodCounter.mlpackage — not committed; train and export it, then drop it into the Watch target |
| Needs to improve | just the final count — a −/+/✓ tap. No tagger, no observer |
| Strength | supervision is nearly free, so it scales with ordinary use |
| Weakness | assumes one dominant tempo per set — it will fight you on drop sets and rest-pause |
A 4-way segmented picker on the watch's idle screen (RecorderModel.RepCountingMode, persisted in
UserDefaults) decides, and it only affects auto-detect sessions — manual sets get their count
from the dial or the phone tagger:
| mode | chain |
|---|---|
| Auto | RepDensityCNN → RepPeriodCNN → RepCounter (first one bundled wins) |
| CNN | RepDensityCNN → RepCounter |
| Period | RepPeriodCNN → RepCounter |
| Unsup | RepCounter only |
Each Core ML counter returns nil when its .mlpackage isn't in the bundle, so the fallback is
automatic and a missing model is never an error — it's just a less accurate count.
All three are scored the same way (exact %, within-1 %, MAE) so they compare head-to-head on your data:
python evaluate_reps.py # unsupervised baseline + per-exercise calibration
python train_reps_windows.py # density model, needs reps.csv
python train_reps_period.py # period model, needs only sets.csvAn honest ceiling. Rep-counting accuracy from a wrist is inherently per-exercise. Arm-dominant lifts (curls, presses, pushdowns) read very cleanly; lower-body work (squats, RDLs, calf raises) gives the wrist a much weaker periodic signal. Expect the former to get good long before the latter, no matter how much data you feed it. Every trainer prints a per-exercise breakdown for exactly this reason.
Every timestamp — samples and set boundaries — comes from the watch's time-since-boot clock
(CMDeviceMotion.timestamp / ProcessInfo.systemUptime), so readings and set intervals are
guaranteed to align without any clock-sync step.
| file | written by | schema |
|---|---|---|
readings.csv |
the watch, every session | subject,session,time_ms,acc_x,acc_y,acc_z,gyro_x,gyro_y,gyro_z |
sets.csv |
the watch, manual sessions | subject,session,exercise,start_ms,end_ms,reps |
detected.csv |
the watch, auto-detect sessions | subject,session,exercise,start_ms,end_ms,reps,confidence,reps_confirmed |
reps.csv |
the phone, live during tagging | subject,session,rep_time_ms,set_start_ms,tap_unix_ms |
The watch writes them as <session-id>_readings.csv etc. and transfers them; the phone files each
one into a date-named folder per session (Documents/exercises/Aug-20/readings.csv, …), matched by
a .sid marker so the second file of a session lands beside the first.
Auto-detect output is kept in a separate file on purpose: an unconfirmed model guess must never
be mistaken for ground truth. A set you mark bad on the watch is written as an interval labeled
discard rather than deleted, so training can drop every window that touches it.
Developer → Build merged export (SessionStore.buildMergedExport) concatenates every received
session into one readings.csv + one sets.csv (+ reps.csv if any taps exist) — headers once,
ready to drop straight into training/data/.
The gate that matters: detected.csv rows are folded into the merged sets.csv only when
reps_confirmed == 1. An uncorrected guess would poison both the calibration fit and the period
model's labels, so it stays out until a human has looked at it.
Everything funnels through two chokepoints — data.load_raw() and rep_events.load_rep_events() —
so the classifier, all three rep counters and the evaluator can never disagree about what the data
says.
load_raw() read + concat every data dir, drop NaNs, sort by (subject, session, time_ms)
└─ apply_exercise_roster() retired exercises → "discard" (default) or "rest"
label_samples() tag each sample with the set interval it falls in, else "rest"
· front trim 0.5 s — absorbs settling after the 3-2-1 countdown
· end trim 1.0 s — absorbs the wind-down before "End Set"
· "discard" intervals applied last, so they win any overlap
make_windows() 100-sample windows, stride 50, majority label
· a window touching a discard interval is dropped
· a window whose top label holds < 80% of it is dropped (too mixed)
make_splits() GroupShuffleSplit BY SUBJECT (≥ 3 subjects), else by window with a warning
compute_norm() per-channel mean/std from TRAIN windows only — baked into the export
Three decisions worth calling out:
- Split by subject, not by window. The biggest real-world error source is a new person. Splitting by window puts a near-duplicate of every validation example into training and reports a beautiful, meaningless number. With fewer than 3 subjects there's no honest way to hold one out, so it splits by window and warns you that the accuracy is optimistic.
- The exercise roster is applied at load. Retiring an exercise on the watch only changes what you
can record next — old sessions still carry it.
RETIRED_POLICY = "drop"throws those windows away rather than calling themrest, because retired exercises are usually mechanically similar to ones you kept (flat vs. incline chest press), and labeling near-identical signal both ways actively damages the class you kept. Put a name back inKEEP_EXERCISESand its data returns — the raw CSVs are never modified. - Old exports can be merged back in. Point
config.EXTRA_DATA_DIRSat any folder with its ownreadings.csv/sets.csv/reps.csv(any subset) — e.g. an export from before a reinstall wiped the session folders that produced it. Each extra directory'ssessionids get a stableextraN:prefix so two unrelated exports can never silently merge into one bout.
Both apps are SwiftUI. The phone app pulls every color, size and radius from DesignSystem.swift
(no literal values in view files) — a dark canvas, one lime accent (#A6F000) reserved for the
thing that matters on each screen, and a cycling tint per exercise.
RecorderModel.Phase drives everything in LiftLoggerWatchApp.swift:
| phase | screen |
|---|---|
.idle |
subject, Start Session, Auto-Detect (AI), the 4-way rep-counter picker, pending-file resend |
.resting |
list of the 15 exercises to start a set, Discard last set, live sets/samples counters, End Session |
.countdown |
3-2-1 in 60 pt rounded digits — "Get into position" |
.inSet |
exercise name, a running timer, End Set |
.enteringReps |
− / count / + dial and Save Set; pre-filled from the phone tagger when it was running (the label turns green to say so) |
.autoDetecting |
live exercise + confidence %, sets logged, the last-set fixer row, End Session |
The fixer row is the heart of the auto-detect screen:
dumbbell curl
− 8 + ✓
LAST SET · TAP TO FIX
Nudging or checking marks the set reps_confirmed. Corrections are buffered in memory and applied
to detected.csv at End Session — never mid-session, because the file handle is still open and
appending; rewriting the file under it would orphan the descriptor and silently drop every set
logged afterward.
| screen | what it's for |
|---|---|
Sessions (SessionsView) |
header with subject pill, a 3-tile summary strip (only SETS takes the accent), the progress module, and one card per received session — title, set count in the exercise tint, exercise chips, and a yellow "Transferring · 1 of 2 files" strip while a transfer is still in flight |
Progress (ProgressModule) |
reps per session for one exercise over the last six sessions, Swift Charts; the exercise list is ordered by how much you actually do it |
Session detail (SessionDetailView) |
the rep rail — one card per exercise, one 44 pt numeral per set on a fixed 3-slot grid, each numeral over a bar whose opacity scales with reps so a card reads as a mini bar chart. Tap a numeral (detected sets only) → Fix Reps. Groups below confidence 0.75 get the uncertain treatment: grey text, dashed bars, a 1 pt outline and a Label button opening the exercise picker |
Developer (DeveloperView) |
off the main flow: subject ID + sync-to-watch, the Rep Tagger, and Build merged export |
Rep Tagger (RepTapView) |
a 280 pt button that arms itself when the watch opens a set. One mutation per press, heavy haptic kept warm, press treatment in a ButtonStyle rather than a gesture — nothing may compete with or swallow a tap, because these timestamps are the ground truth |
Nothing is thrown away, and nothing counts as truth until you've looked at it.
On the wrist, during an auto-detect session — tap −/+ to nudge the last set's count, or ✓
if the guess was already right. Either marks it confirmed.
On the phone — open a session, tap any detected set's rep numeral for the Fix Reps sheet, or
tap Label on an uncertain card to name a misidentified exercise. SessionStore.confirmReps
rewrites that row's reps and stamps reps_confirmed=1; confirmLabel rewrites the exercise at
confidence 1.0. Both edit detected.csv in place, so corrections survive into the export rather
than living only in the UI.
Then Build merged export folds only confirmed rows into sets.csv, which is what
evaluate_reps.py's calibration and train_reps_period.py's training both read.
- Xcode 26 or newer, macOS host
- An Apple Watch paired to an iPhone (deployment targets: watchOS 26.5, iOS 26.5). The simulator is fine for UI work but produces no IMU, so it can't record.
- Python 3.9+ for the training pipeline (only if you want to retrain)
git clone <this-repo>
cd Watch-Exercise-Tracker
open LiftLogger.xcodeproj- Select the LiftLogger scheme → your iPhone → Run.
- Select the LiftLogger Watch App scheme → your watch → Run.
- On the phone, open ⚙︎ Developer, set a subject ID (e.g.
S01), tap Sync subject to watch. - On the watch, tap Auto-Detect (AI) and lift.
Notes:
- Signing — set your own team on both targets; the bundle IDs (
com.ryandong.LiftLogger*) are placeholders. - A free Apple ID is enough. HealthKit is deliberately not used — it's a restricted entitlement
requiring a paid Developer Program membership — so the watch keeps the session alive without an
HKWorkoutSession, and the entitlements file is empty. - Adding a model is a file copy. The project uses Xcode's file-system-synchronized groups, so
dropping a
.mlpackageintoLiftLogger Watch App/adds it to the target — no drag-into-Xcode dance, noproject.pbxprojedit. - Committed today:
LiftLoggerClassifier.mlpackageandLiftLoggerRepCounter.mlpackage.LiftLoggerRepPeriodCounter.mlpackageis not — train it if you want Period mode.
cd training
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt # numpy, pandas, scikit-learn, torch, coremltoolsNothing to collect yet? Exercise the whole pipeline on fake data first:
python make_synthetic.py # writes data/readings.csv + data/sets.csv in the real schema
python train.py # → artifacts/cnn.pt, classes.json, norm.npz, metrics.txt
python export_coreml.py # → artifacts/LiftLoggerClassifier.mlpackageThe synthetic exercises are distinct sine patterns, so accuracy will be near-perfect. That proves the plumbing, nothing else.
- Record sessions on the watch; they sync to the phone on their own.
- Developer → Build merged export → AirDrop / Save
readings.csv,sets.csv(andreps.csv). - Drop them into
training/data/. - Train, export, copy the
.mlpackageintoLiftLogger Watch App/, rebuild the watch app.
Run everything from inside training/ with the venv active.
# exercise classifier
python train.py # subject-held-out; prints per-class F1 + confusion matrix
python baseline_rf.py # the Random Forest number the CNN must beat
python export_coreml.py # → LiftLoggerClassifier.mlpackage
# rep counting — compare all three on the same data
python evaluate_reps.py # unsupervised + fits artifacts/rep_calibration.json
python train_reps_windows.py # density model (needs reps.csv taps) ← preferred
python train_reps_model.py # legacy whole-bout density baseline
python train_reps_period.py # period model (needs only sets.csv reps)
# export a rep counter
python export_reps_coreml.py # windowed density → LiftLoggerRepCounter.mlpackage
python export_reps_coreml.py --bout # legacy bout density
python export_reps_coreml.py --period # period model → LiftLoggerRepPeriodCounter.mlpackage
# hands-off: watch a folder and retrain whenever the CSVs change
python auto_retrain.py # poll every 30 s
python auto_retrain.py --once
python auto_retrain.py --watch-dir ~/Documents/LiftLoggerExports # e.g. an iCloud folderWhat to look at, in order: the per-class F1 in metrics.txt (a high overall accuracy hides a
class the model never gets right), then the confusion matrix (mechanically similar moves — bicep vs
hammer curl — blur together and tell you where to collect more or merge classes), then the
per-exercise rep MAE.
What actually moves accuracy: data variety over model complexity. Multiple subjects, both
wrists, the watch rotated differently, varied tempo and load. Tens of sets per exercise per person.
Class weights are already applied during training, because rest will otherwise dominate
every batch — and metrics.txt reports per-class F1 so it can't hide behind overall accuracy.
Every tunable lives in training/config.py with the reasoning next to it — it's the best single
file to read if you want to understand the system.
| path | what it is |
|---|---|
LiftLogger Watch App/ |
the watch app — RecorderModel.swift (sensor loop, live inference, session logging, on-wrist correction, RepCounter + RepDensityCounter), RepPeriodCounter.swift, LiftLoggerWatchApp.swift (all six phase screens), bundled .mlpackages |
LiftLogger/ |
the iPhone app — SessionStore.swift (receives files, corrections, merged export, live tagging), SessionsView.swift, SessionDetailView.swift (rep rail, Fix Reps, label picker), SessionSummary.swift (CSV → exercise/set model), ProgressModule.swift, RepTapView.swift, DesignSystem.swift |
training/ |
the PyTorch pipeline — see training/README.md for the full model-choice writeup |
docs/ |
the logo and the animated demos above (make_demo_svgs.py regenerates them) |
design.md |
the phone app's design spec — every § reference in the iOS view files points here |
LiftLogger.xcodeproj |
both targets, plus the XCTest/UI-test stubs |
liftlogger files/ |
the original standalone data-collection version, superseded by the two app folders above — kept for reference |
liftlogger-course.html |
a long-form write-up of how the whole thing was built, as a self-contained page |
- One wrist, one watch. Everything assumes the watch is on the working arm's wrist.
- Lower-body accuracy is capped by physics, not by the model — see the note in Which one runs.
- The bundled classifier covers 7 exercises + rest, not the full 15-exercise roster the UI offers. Retrain to close the gap.
- Single dominant tempo per set is baked into all three rep counters; drop sets and rest-pause sets will be counted as one bout at one tempo.
reps.csvneeds a second person (or a very patient you) to tap along — which is exactly the constraint the period model exists to remove.
Most of this repo — the Swift for both apps, the PyTorch training pipeline, the Core ML export scripts, and this README — was written with Claude Code, Anthropic's agentic coding tool, working from my direction. I chose the problem, the sensor and model architecture, and the three rep-counting approaches; I collected the training data, ran the training, and tested everything on a real wrist. But the line-by-line implementation is largely Claude's, and it would be dishonest to present it otherwise.
Two things follow from that, and they're worth stating plainly:
- The design decisions are real, and so are the results. The numbers in The models come from actual training runs on actual recorded sets, not from anything a model asserted about its own output.
- Read the code before you trust it. Agent-written code carries agent-written mistakes, and I have not hand-audited every path. The Limitations above are the ones I know about.