Dataset access came through on Monday. The first thing I did was not train anything.
25 subjects. 200 sequences. 48,398 frames. Before any of that touches a model, I wanted to know what was actually in it.
The thing hiding in the filenames
Every sequence name in UNBC-McMaster parses the same way: subject, trial, movement, limb. I'd read that structure as bookkeeping and moved on the first time I saw it.
The limb field is aff or unaff, and unaff means the unaffected shoulder: the clinician ran the same range-of-motion test on the arm that doesn't hurt. 109 sequences are the painful arm. 91 are the other one.
That is a control condition, built into the dataset, that none of us had noticed. We had been thinking of our negatives as "frames that happened to be neutral," background we'd absorb wherever it showed up, when closer to a third of the data is structurally non-painful by design, from a limb a clinician chose specifically because it wasn't the problem. It also hands us a test nobody had proposed: a model that understands pain should stay quiet on unaff sequences almost by construction. If it doesn't, that's not noise to average out, that's the model reacting to the movement itself rather than the pain.
The question that mattered
Our whole architecture is a bet that pain has temporal structure. The LSTM reads a 32-frame window and outputs one confidence number for it. If PSPI (the frame-level pain intensity score the dataset ships with) spiked for two or three frames and vanished, a 32-frame window would have almost nothing to learn from, and we would have built the wrong shape of model before writing a line of training code.
So before committing to that architecture, I plotted it.
It builds. Every one of the six highest-pain sequences shows the same shape: a rise over 20 to 40 frames, a plateau, a decay, not a spike. That validates the window size we'd already committed to, and it means max-in-window labelling (taking the highest PSPI value inside a window as that window's label) is a defensible choice rather than a shortcut we'd have to justify later.
What is harder
Most frames are neutral. On the linear axis almost everything else disappears into the floor; the log axis is the only reason you can see that pain exists in this dataset at all. The imbalance is structural. It is what pain looks like on video, not a preprocessing problem, and no amount of augmentation manufactures pain that was not there in the first place.
Pick threshold 1 and a quarter of our windows are positive. Pick threshold 5 and it's under 5%. That is a modelling choice with real consequences for precision, recall, and what an "alert" ends up meaning at the showcase, and we have not settled it.
The uncomfortable one
I want to be honest about this one rather than let it sit quietly in a supplementary folder. VAS is what the patient said their pain was. PSPI is what a trained coding system read from their face. The correlation between them is r = 0.59. Against a trained observer's rating it rises to 0.74: a face is easier to grade than it is to match against what the patient says they are feeling, and the two only agree moderately even when a human is doing the reading. That is the ceiling on what any camera-based system can do, including ours, and it exists before a single neural network gets involved. This belongs on the project page, next to the numbers we're claiming, not buried at the bottom of a journal post where it's easy to skim past.
Close unresolved
25 subjects means a single train/test split is fragile. Hold out five people and the result you report depends materially on which five: a different split can move our headline number more than a genuine architecture change would. We are switching to grouped k-fold, splitting by subject so no person's frames appear in both train and test, and reporting a range rather than a point estimate.
The threshold is still open. I don't have a principled answer yet for which PSPI cutoff defines "positive," and I would rather say that plainly than pick one now and back-justify it later.