Skip to content

Repository files navigation

LiftLogger

LiftLogger

Your Apple Watch already knows which lift you're doing. This teaches it to write it down.

Automatic exercise recognition and rep counting from wrist IMU, running entirely on-device — with a companion iPhone app where every correction you make becomes tomorrow's training data.

watchOS 26.5+ iOS 26.5+ Swift 5.0 SwiftUI · Core ML PyTorch 100% on-device

How it works · The models · Data · The UI · Install · Training


Demo

LiftLogger on Apple Watch: pick a rep counter, auto-detect a curl, fix the count on the wrist

Watch — pick a rep counter, tap Auto-Detect, lift. The model names the exercise live, closes the set when you stop, and offers the count for a one-tap fix.

LiftLogger on iPhone: sessions list, the rep rail, and the Fix Reps sheet

iPhone — sessions arrive on their own. Tap a session for the rep rail, tap any rep numeral to correct it, and that row is promoted to ground truth.

Both demos are animated SVG mockups built from the real view code and design tokens (docs/make_demo_svgs.py), not screen recordings — the layout, copy and palette are lifted from LiftLoggerWatchApp.swift, SessionsView.swift, SessionDetailView.swift and DesignSystem.swift. To swap in a real capture, drop a .gif in docs/ and change the <img src> above.


Contents


What it does

You start a session on the watch and lift. That's the whole interaction.

  • No exercise selection. A 1-D CNN reads the wrist IMU and names the movement ~2× per second.
  • No rep counting. When a set closes, one of three rep models counts it from the same signal.
  • No manual sync. WatchConnectivity ships the session's CSVs to the phone at End Session.
  • No wasted correction. Fixing a wrong count on the wrist or the phone marks that set reps_confirmed, and confirmed rows are exactly what the next training run reads.
  • Nothing leaves the devices. Core ML on the Neural Engine; no server, no account, no network calls.

How it works

On the watch — from wrist motion to a logged set:

flowchart TD
    A["CMMotionManager · 50 Hz<br/>acc_xyz + gyro_xyz"] --> B["rolling 2 s window<br/>100 samples, re-run every 25"]
    B --> C{"CNN1D classifier<br/>~2 predictions / sec"}
    C -->|"confidence below 0.6,<br/>or the class is still unstable"| B
    C -->|"same class 3 windows running"| D["SET OPENS<br/>buffer the bout + 3 s of lead-in"]
    D --> E{"rest stable 5 windows<br/>and 5 s grace elapsed?"}
    E -->|no| D
    E -->|yes| F["SET CLOSES · hand the bout to a rep counter"]
    F --> G["RepCounter<br/>autocorrelation"]
    F --> H["RepDensityCNN<br/>density curve"]
    F --> I["RepPeriodCNN<br/>rep period"]
    G & H & I --> J["detected.csv · reps_confirmed = 0"]
Loading

Off the watch — the flywheel that makes the next model better:

flowchart LR
    A["detected.csv<br/>reps_confirmed = 0"] -->|WatchConnectivity| B["iPhone · Sessions"]
    B --> C["you fix or confirm<br/>−/+/✓ on the wrist,<br/>Fix Reps on the phone"]
    C --> D["reps_confirmed = 1"]
    D --> E["Build merged export<br/>readings · sets · reps"]
    E --> F["training/<br/>PyTorch"]
    F --> G[".mlpackage"]
    G -->|"drop into the Watch target"| H["a better model<br/>next session"]
Loading

The loop closes: the sets you correct are the sets the next model learns from.

Every constant that has to agree across the Swift and Python sides is named in both files with a MUST match comment — WINDOW/windowSize, FS/fs, REP_WIN_STRIDE/repWinStride, REST_LABEL/restLabel, DISCARD_LABEL/discardLabel.

Live detection tuning

Every knob below lives at the top of RecorderModel.swift with the reasoning attached.

knob value why
sample rate 50 Hz matches config.FS; fast enough for a 0.5 s rep, cheap enough to run all session
inference window 100 samples (2 s) matches config.WINDOW — a couple of reps of context
inference stride 25 samples a prediction ~2× per second
confidence floor 0.60 below this the window counts as rest; don't log guesses
windows to open a set 3 debounce, so one lucky window can't start a set
windows to close a set 5 deliberately harder to leave an exercise than to enter one
rest grace 5 s a hold at the top, a breath, a grip reset all look like rest to a 2 s window — this keeps the set from fragmenting
minimum set 1.5 s anything shorter is noise, not a set
bout lead-in 150 samples (3 s) kept before the set commits so the first reps aren't clipped
bout cap 6000 samples (120 s) hard memory ceiling per set
countdown 3 s manual sessions only — time to get into position before start_ms is stamped

The models

Four models total: one that answers what you're doing, and three competing answers to how many. All of them are the same shape of thing — small 1-D convnets over 6-channel IMU — and all of them convert cleanly to Core ML and run on the Watch's Neural Engine.

1. Exercise classifier — CNN1D

training/models.py · trained by train.py · exported by export_coreml.py

input   (1, 100, 6)   raw acc_x/y/z, gyro_x/y/z @ 50 Hz — the model normalizes internally
        conv7 → 64  → BN → ReLU → maxpool2
        conv5 → 128 → BN → ReLU → maxpool2
        conv3 → 128 → BN → ReLU → global average pool
        dropout 0.3 → linear
output  softmax over N exercises + "rest"          ~150k parameters

Why a temporal ConvNet and not something else:

  • Translation invariance. A 2 s window can start anywhere in a rep — at the top, mid-descent, in the pause. Convolution + global average pooling means the model doesn't care.
  • Small-data robustness. Global average pooling instead of a flattened dense layer keeps the parameter count at ~150k, which matters when a class has 40 examples, not 40,000.
  • It deploys. No recurrence, no attention, no per-prediction search over a training set — a clean torch.jit.trace → coremltools conversion that the Neural Engine likes.

The normalization statistics are computed on training windows only and baked into the exported graph (ExportWrapper), so the Swift side hands over a raw sensor buffer and reads back probabilities. There is no feature engineering to keep in sync across two languages.

A Random Forest on hand-crafted window features (baseline_rf.py, features.py) is kept around as the number to beat. If the CNN can't clearly beat it, the bottleneck is data, not architecture. training/README.md has the full ranking of alternatives considered (DeepConvLSTM, TCN, k-NN+DTW, transformers) and why each was accepted or rejected for a watch.

The 15-exercise roster (RecorderModel.exercises, mirrored by config.KEEP_EXERCISES): incline chest press · machine chest press · machine shoulder press · wide-grip machine row · cable push down · overhead triceps · dumbbell hammer curl · dumbbell curl · cable curl · forearm raises · lat pulldown · squat · dumbbell RDL · machine calf raise · dumbbell Bulgarian split squat — plus an implicit rest class for everything else. The phone's label picker (SessionStore.knownExercises) additionally lists a few retired names, so old sessions recorded before the roster changed can still be relabeled.

Note on the bundled checkpoint. The .mlpackage committed in LiftLogger Watch App/ was exported when the roster was 7 exercises + rest (cable_push_down, dumbbell_hammer_curl, forearm_raises, incline_chest_press, machine_row_wide, machine_shoulder_press, overhead_triceps). The watch UI and config.KEEP_EXERCISES now offer 15 — retrain and re-export to cover the rest.

2. Rep counter A — unsupervised periodicity (RepCounter)

training/reps.py ⟷ RepCounter in RecorderModel.swift · evaluated by evaluate_reps.py

A rep is one cycle of a quasi-periodic movement, so no learning is required at all:

  1. For each of the 6 channels, compute the normalized autocorrelation over lags in 0.5–4.0 s (REP_PERIOD_RANGE_S, i.e. 0.25–2 reps/sec).
  2. Keep the channel with the strongest peak — the axis that best expresses this exercise picks itself, which is why one implementation covers curls and squats alike.
  3. Its peak lag is the rep period T; reps ≈ bout_length / T.
  4. Cross-check with a peak count on that channel (minimum spacing 0.6·T, prominence 0.30 × signal std). If autocorrelation is weak (< 0.25), trust the peaks instead; if the two agree within 1, prefer the peaks — they handle partial first/last reps slightly better.
Needs to run nothing — no weights, no bundle, always available
Needs to improve nothing trainable; evaluate_reps.py fits a per-exercise linear correction true ≈ a·pred + b from confirmed counts (≥ 5 sets per exercise) and writes artifacts/rep_calibration.json
Strength zero cold-start, fully interpretable, identical logic in Python and Swift
Weakness a systematic bias per exercise, and it degrades where the wrist barely moves periodically

This is the floor every learned counter has to beat, and the fallback at the end of every chain.

3. Rep counter B — CNN density curve (RepDensityCNN)

training/models.py · data by rep_windows.py · trained by train_reps_windows.py · exported by export_reps_coreml.py ⟷ RepDensityCounter in RecorderModel.swift

Instead of regressing a count, this predicts where the reps are and integrates:

input   (1, 400, 6)   an 8 s window at real 50 Hz
        conv7 → 64 stem, then 7 dilated conv3 blocks, dilations (1,2,4,8,16,32,64)
        NO pooling over time — output length stays 400
        conv1x1 → Softplus
output  (1, 400)      per-frame rep density ≥ 0;  reps = Σ density

The labels are per-rep tap timestamps from the phone's Rep Tagger (RepTapView): an observer taps once per rep while you lift, each tap becomes a unit-area Gaussian (σ = 0.20 s) on the density curve, so the curve integrates back to the true count. On the watch, RepDensityCounter slides the same 8 s window across the bout at a 1 s stride and overlap-adds by averaging — dividing by how many windows covered each frame, so a rep in an overlap region isn't counted twice — exactly mirroring rep_windows.bout_count in Python.

Three design details that are easy to get wrong, all handled explicitly:

  • Receptive field must span a rep. RF = 7 + 2·Σdilations. The legacy whole-bout dilations (1,2,4,8) give 37 frames — fine when a frame is 1/256th of a set, but only 0.74 s in real time, less than one rep of a slow lift. REP_WIN_DILATIONS reaches 261 frames ≈ 5.2 s.
  • Head bias initialization. A real density averages ≈ 0.01 reps/frame, while an untuned Softplus head starts at 0.69 — 70× too high. The trainer initializes the output bias to inverse_softplus(mean target). Without it, the model burns its entire epoch budget just deflating, which looks exactly like convergence.
  • Windows, not whole bouts. An earlier path (train_reps_model.py, kept for comparison) resampled each set to 256 frames — which made a 10 s set and a 50 s set at the same tempo look 5× different, and gave ~50 training examples where the classifier had thousands. Fixed-second windows turn one 30 s set into ~23 examples and keep tempo in real Hz.
Needs to run LiftLoggerRepCounter.mlpackage bundled (it is — the windowed variant)
Needs to improve per-rep tap timestamps in reps.csv — dense, accurate, effortful: someone has to tap along
Strength most accurate when it has data, and the only counter that tells you where each rep was
Weakness the tapping bottleneck; every new exercise needs a live tagger session

4. Rep counter C — CNN period regression (RepPeriodCNN)

training/models.py · trained by train_reps_period.py · exported by export_reps_coreml.py --period ⟷ RepPeriodCounter.swift

The counter built specifically to remove the tapping bottleneck. It trains on nothing but the final rep integer already sitting in sets.csv:

input   (1, 600, 6)   12 s of the bout, edge-padded if short / cropped from the start if long
        same conv/pool trunk as CNN1D (global average pooled)
        dropout → linear → 1 scalar
output  log(period_seconds)  →  exported wrapper returns exp(·), i.e. SECONDS
                                the watch computes  reps = round(bout_duration / period)

Two choices carry this model:

  • Predict the period, not the count. A count entangles tempo with however long the set happened to run; the period is duration-independent and lives in a narrow physical band (0.5–4.0 s). That makes a single scalar per set enough supervision to actually learn from — a far more sample-efficient regression target.
  • Train in log-space. Periods are ratio-scale: a 2 s rep vs a 1 s rep is "twice as slow", not "1 s slower". Log targets stop fast and slow reps from contributing wildly different-magnitude losses.

Its training set grows every time you use the app: every hand-dialed manual set, plus every auto-detected set you nudged or checked off — on the wrist or in the phone's Fix Reps sheet. Sets with fewer than 2 reps are ignored (REP_PERIOD_MIN_REPS).

Needs to run LiftLoggerRepPeriodCounter.mlpackage — not committed; train and export it, then drop it into the Watch target
Needs to improve just the final count — a −/+/✓ tap. No tagger, no observer
Strength supervision is nearly free, so it scales with ordinary use
Weakness assumes one dominant tempo per set — it will fight you on drop sets and rest-pause

Which one runs

A 4-way segmented picker on the watch's idle screen (RecorderModel.RepCountingMode, persisted in UserDefaults) decides, and it only affects auto-detect sessions — manual sets get their count from the dial or the phone tagger:

mode chain
Auto RepDensityCNN → RepPeriodCNN → RepCounter (first one bundled wins)
CNN RepDensityCNN → RepCounter
Period RepPeriodCNN → RepCounter
Unsup RepCounter only

Each Core ML counter returns nil when its .mlpackage isn't in the bundle, so the fallback is automatic and a missing model is never an error — it's just a less accurate count.

All three are scored the same way (exact %, within-1 %, MAE) so they compare head-to-head on your data:

python evaluate_reps.py        # unsupervised baseline + per-exercise calibration
python train_reps_windows.py   # density model, needs reps.csv
python train_reps_period.py    # period model, needs only sets.csv

An honest ceiling. Rep-counting accuracy from a wrist is inherently per-exercise. Arm-dominant lifts (curls, presses, pushdowns) read very cleanly; lower-body work (squats, RDLs, calf raises) gives the wrist a much weaker periodic signal. Expect the former to get good long before the latter, no matter how much data you feed it. Every trainer prints a per-exercise breakdown for exactly this reason.


Data & how it is loaded

What the watch writes

Every timestamp — samples and set boundaries — comes from the watch's time-since-boot clock (CMDeviceMotion.timestamp / ProcessInfo.systemUptime), so readings and set intervals are guaranteed to align without any clock-sync step.

file written by schema
readings.csv the watch, every session subject,session,time_ms,acc_x,acc_y,acc_z,gyro_x,gyro_y,gyro_z
sets.csv the watch, manual sessions subject,session,exercise,start_ms,end_ms,reps
detected.csv the watch, auto-detect sessions subject,session,exercise,start_ms,end_ms,reps,confidence,reps_confirmed
reps.csv the phone, live during tagging subject,session,rep_time_ms,set_start_ms,tap_unix_ms

The watch writes them as <session-id>_readings.csv etc. and transfers them; the phone files each one into a date-named folder per session (Documents/exercises/Aug-20/readings.csv, …), matched by a .sid marker so the second file of a session lands beside the first.

Auto-detect output is kept in a separate file on purpose: an unconfirmed model guess must never be mistaken for ground truth. A set you mark bad on the watch is written as an interval labeled discard rather than deleted, so training can drop every window that touches it.

What the phone merges

Developer → Build merged export (SessionStore.buildMergedExport) concatenates every received session into one readings.csv + one sets.csv (+ reps.csv if any taps exist) — headers once, ready to drop straight into training/data/.

The gate that matters: detected.csv rows are folded into the merged sets.csv only when reps_confirmed == 1. An uncorrected guess would poison both the calibration fit and the period model's labels, so it stays out until a human has looked at it.

How the trainer loads it

Everything funnels through two chokepoints — data.load_raw() and rep_events.load_rep_events() — so the classifier, all three rep counters and the evaluator can never disagree about what the data says.

load_raw()            read + concat every data dir, drop NaNs, sort by (subject, session, time_ms)
  └─ apply_exercise_roster()   retired exercises → "discard" (default) or "rest"
label_samples()       tag each sample with the set interval it falls in, else "rest"
                        · front trim 0.5 s  — absorbs settling after the 3-2-1 countdown
                        · end trim   1.0 s  — absorbs the wind-down before "End Set"
                        · "discard" intervals applied last, so they win any overlap
make_windows()        100-sample windows, stride 50, majority label
                        · a window touching a discard interval is dropped
                        · a window whose top label holds < 80% of it is dropped (too mixed)
make_splits()         GroupShuffleSplit BY SUBJECT (≥ 3 subjects), else by window with a warning
compute_norm()        per-channel mean/std from TRAIN windows only — baked into the export

Three decisions worth calling out:

  • Split by subject, not by window. The biggest real-world error source is a new person. Splitting by window puts a near-duplicate of every validation example into training and reports a beautiful, meaningless number. With fewer than 3 subjects there's no honest way to hold one out, so it splits by window and warns you that the accuracy is optimistic.
  • The exercise roster is applied at load. Retiring an exercise on the watch only changes what you can record next — old sessions still carry it. RETIRED_POLICY = "drop" throws those windows away rather than calling them rest, because retired exercises are usually mechanically similar to ones you kept (flat vs. incline chest press), and labeling near-identical signal both ways actively damages the class you kept. Put a name back in KEEP_EXERCISES and its data returns — the raw CSVs are never modified.
  • Old exports can be merged back in. Point config.EXTRA_DATA_DIRS at any folder with its own readings.csv / sets.csv / reps.csv (any subset) — e.g. an export from before a reinstall wiped the session folders that produced it. Each extra directory's session ids get a stable extraN: prefix so two unrelated exports can never silently merge into one bout.

The UI

Both apps are SwiftUI. The phone app pulls every color, size and radius from DesignSystem.swift (no literal values in view files) — a dark canvas, one lime accent (#A6F000) reserved for the thing that matters on each screen, and a cycling tint per exercise.

Watch — six phases, one state machine

RecorderModel.Phase drives everything in LiftLoggerWatchApp.swift:

phase screen
.idle subject, Start Session, Auto-Detect (AI), the 4-way rep-counter picker, pending-file resend
.resting list of the 15 exercises to start a set, Discard last set, live sets/samples counters, End Session
.countdown 3-2-1 in 60 pt rounded digits — "Get into position"
.inSet exercise name, a running timer, End Set
.enteringReps − / count / + dial and Save Set; pre-filled from the phone tagger when it was running (the label turns green to say so)
.autoDetecting live exercise + confidence %, sets logged, the last-set fixer row, End Session

The fixer row is the heart of the auto-detect screen:

   dumbbell curl
   −   8   +   ✓
LAST SET · TAP TO FIX

Nudging or checking marks the set reps_confirmed. Corrections are buffered in memory and applied to detected.csv at End Session — never mid-session, because the file handle is still open and appending; rewriting the file under it would orphan the descriptor and silently drop every set logged afterward.

iPhone — four screens and a chart

screen what it's for
Sessions (SessionsView) header with subject pill, a 3-tile summary strip (only SETS takes the accent), the progress module, and one card per received session — title, set count in the exercise tint, exercise chips, and a yellow "Transferring · 1 of 2 files" strip while a transfer is still in flight
Progress (ProgressModule) reps per session for one exercise over the last six sessions, Swift Charts; the exercise list is ordered by how much you actually do it
Session detail (SessionDetailView) the rep rail — one card per exercise, one 44 pt numeral per set on a fixed 3-slot grid, each numeral over a bar whose opacity scales with reps so a card reads as a mini bar chart. Tap a numeral (detected sets only) → Fix Reps. Groups below confidence 0.75 get the uncertain treatment: grey text, dashed bars, a 1 pt outline and a Label button opening the exercise picker
Developer (DeveloperView) off the main flow: subject ID + sync-to-watch, the Rep Tagger, and Build merged export
Rep Tagger (RepTapView) a 280 pt button that arms itself when the watch opens a set. One mutation per press, heavy haptic kept warm, press treatment in a ButtonStyle rather than a gesture — nothing may compete with or swallow a tap, because these timestamps are the ground truth

The correction loop

Nothing is thrown away, and nothing counts as truth until you've looked at it.

On the wrist, during an auto-detect session — tap −/+ to nudge the last set's count, or ✓ if the guess was already right. Either marks it confirmed.

On the phone — open a session, tap any detected set's rep numeral for the Fix Reps sheet, or tap Label on an uncertain card to name a misidentified exercise. SessionStore.confirmReps rewrites that row's reps and stamps reps_confirmed=1; confirmLabel rewrites the exercise at confidence 1.0. Both edit detected.csv in place, so corrections survive into the export rather than living only in the UI.

Then Build merged export folds only confirmed rows into sets.csv, which is what evaluate_reps.py's calibration and train_reps_period.py's training both read.


Installation

Prerequisites

  • Xcode 26 or newer, macOS host
  • An Apple Watch paired to an iPhone (deployment targets: watchOS 26.5, iOS 26.5). The simulator is fine for UI work but produces no IMU, so it can't record.
  • Python 3.9+ for the training pipeline (only if you want to retrain)

The apps

git clone <this-repo>
cd Watch-Exercise-Tracker
open LiftLogger.xcodeproj
  1. Select the LiftLogger scheme → your iPhone → Run.
  2. Select the LiftLogger Watch App scheme → your watch → Run.
  3. On the phone, open ⚙︎ Developer, set a subject ID (e.g. S01), tap Sync subject to watch.
  4. On the watch, tap Auto-Detect (AI) and lift.

Notes:

  • Signing — set your own team on both targets; the bundle IDs (com.ryandong.LiftLogger*) are placeholders.
  • A free Apple ID is enough. HealthKit is deliberately not used — it's a restricted entitlement requiring a paid Developer Program membership — so the watch keeps the session alive without an HKWorkoutSession, and the entitlements file is empty.
  • Adding a model is a file copy. The project uses Xcode's file-system-synchronized groups, so dropping a .mlpackage into LiftLogger Watch App/ adds it to the target — no drag-into-Xcode dance, no project.pbxproj edit.
  • Committed today: LiftLoggerClassifier.mlpackage and LiftLoggerRepCounter.mlpackage. LiftLoggerRepPeriodCounter.mlpackage is not — train it if you want Period mode.

The training pipeline

cd training
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt        # numpy, pandas, scikit-learn, torch, coremltools

Nothing to collect yet? Exercise the whole pipeline on fake data first:

python make_synthetic.py    # writes data/readings.csv + data/sets.csv in the real schema
python train.py             # → artifacts/cnn.pt, classes.json, norm.npz, metrics.txt
python export_coreml.py     # → artifacts/LiftLoggerClassifier.mlpackage

The synthetic exercises are distinct sine patterns, so accuracy will be near-perfect. That proves the plumbing, nothing else.

With real data

  1. Record sessions on the watch; they sync to the phone on their own.
  2. Developer → Build merged export → AirDrop / Save readings.csv, sets.csv (and reps.csv).
  3. Drop them into training/data/.
  4. Train, export, copy the .mlpackage into LiftLogger Watch App/, rebuild the watch app.

Training recipes

Run everything from inside training/ with the venv active.

# exercise classifier
python train.py                      # subject-held-out; prints per-class F1 + confusion matrix
python baseline_rf.py                # the Random Forest number the CNN must beat
python export_coreml.py              # → LiftLoggerClassifier.mlpackage

# rep counting — compare all three on the same data
python evaluate_reps.py              # unsupervised + fits artifacts/rep_calibration.json
python train_reps_windows.py         # density model (needs reps.csv taps)   ← preferred
python train_reps_model.py           # legacy whole-bout density baseline
python train_reps_period.py          # period model (needs only sets.csv reps)

# export a rep counter
python export_reps_coreml.py             # windowed density  → LiftLoggerRepCounter.mlpackage
python export_reps_coreml.py --bout      # legacy bout density
python export_reps_coreml.py --period    # period model      → LiftLoggerRepPeriodCounter.mlpackage

# hands-off: watch a folder and retrain whenever the CSVs change
python auto_retrain.py                   # poll every 30 s
python auto_retrain.py --once
python auto_retrain.py --watch-dir ~/Documents/LiftLoggerExports   # e.g. an iCloud folder

What to look at, in order: the per-class F1 in metrics.txt (a high overall accuracy hides a class the model never gets right), then the confusion matrix (mechanically similar moves — bicep vs hammer curl — blur together and tell you where to collect more or merge classes), then the per-exercise rep MAE.

What actually moves accuracy: data variety over model complexity. Multiple subjects, both wrists, the watch rotated differently, varied tempo and load. Tens of sets per exercise per person. Class weights are already applied during training, because rest will otherwise dominate every batch — and metrics.txt reports per-class F1 so it can't hide behind overall accuracy.

Every tunable lives in training/config.py with the reasoning next to it — it's the best single file to read if you want to understand the system.


Repo layout

path what it is
LiftLogger Watch App/ the watch app — RecorderModel.swift (sensor loop, live inference, session logging, on-wrist correction, RepCounter + RepDensityCounter), RepPeriodCounter.swift, LiftLoggerWatchApp.swift (all six phase screens), bundled .mlpackages
LiftLogger/ the iPhone app — SessionStore.swift (receives files, corrections, merged export, live tagging), SessionsView.swift, SessionDetailView.swift (rep rail, Fix Reps, label picker), SessionSummary.swift (CSV → exercise/set model), ProgressModule.swift, RepTapView.swift, DesignSystem.swift
training/ the PyTorch pipeline — see training/README.md for the full model-choice writeup
docs/ the logo and the animated demos above (make_demo_svgs.py regenerates them)
design.md the phone app's design spec — every § reference in the iOS view files points here
LiftLogger.xcodeproj both targets, plus the XCTest/UI-test stubs
liftlogger files/ the original standalone data-collection version, superseded by the two app folders above — kept for reference
liftlogger-course.html a long-form write-up of how the whole thing was built, as a self-contained page

Limitations

  • One wrist, one watch. Everything assumes the watch is on the working arm's wrist.
  • Lower-body accuracy is capped by physics, not by the model — see the note in Which one runs.
  • The bundled classifier covers 7 exercises + rest, not the full 15-exercise roster the UI offers. Retrain to close the gap.
  • Single dominant tempo per set is baked into all three rep counters; drop sets and rest-pause sets will be counted as one bout at one tempo.
  • reps.csv needs a second person (or a very patient you) to tap along — which is exactly the constraint the period model exists to remove.

How this was built

Most of this repo — the Swift for both apps, the PyTorch training pipeline, the Core ML export scripts, and this README — was written with Claude Code, Anthropic's agentic coding tool, working from my direction. I chose the problem, the sensor and model architecture, and the three rep-counting approaches; I collected the training data, ran the training, and tested everything on a real wrist. But the line-by-line implementation is largely Claude's, and it would be dishonest to present it otherwise.

Two things follow from that, and they're worth stating plainly:

  • The design decisions are real, and so are the results. The numbers in The models come from actual training runs on actual recorded sets, not from anything a model asserted about its own output.
  • Read the code before you trust it. Agent-written code carries agent-written mistakes, and I have not hand-audited every path. The Limitations above are the ones I know about.
Built with Claude Code. Everything runs on-device.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages