For about ten days in week five, none of our models were ready to run. So we wrote the entire inference loop anyway, against three stubs that returned random numbers of the right shape.

It felt like procrastination. It was the most useful ten days of the project so far.

The stubs

# pulse/stubs.py: stand-ins with the right shapes and nothing else
def detect_face(frame):
    # 1 in 8 frames has no face, like a real room
    if random.random() < 0.125: return None
    return np.zeros((224, 224, 3), np.uint8)

def features(crop): return np.random.randn(256).astype(np.float32)

def confidence(window): return float(np.clip(np.random.normal(0.4, 0.2), 0, 1))

FIGURE 1 · THREE FUNCTIONS THAT KNOW NOTHING AND STILL FOUND FOUR BUGS.

The one decision worth copying is the 1 in 8 frames has no face line. A stub that always succeeds tests nothing. A stub that fails at a realistic rate turns the missing-face path from a hypothetical into something you trip over on the first run.

Screenshot of a Colab notebook cell titled Transforms, defining train_transform with RandomHorizontalFlip, RandomRotation of 10 degrees, ColorJitter and Normalize, and a val_transform that only normalises.
FIGURE 2 · THE PREPROCESSING THE REAL MODELS WILL SLOT INTO: AUGMENTATION ON TRAIN, NORMALISATION ONLY ON VALIDATION.

What it caught

  • The buffer never drained. On a missing face we skipped the frame but also skipped the buffer eviction, so the window silently held frames from ten seconds earlier.
  • Alerts fired every frame. Above threshold, we sent a notification per frame, thirty per second. We now debounce to one alert per 30-second window unless confidence rises materially.
  • Timestamps came from the wrong clock. The alert carried the time the notification was built, not the time of the frame that triggered it. Off by up to 400ms, which sounds small until a carer asks what happened when.
  • Startup lied. For the first 32 frames the window is not full, and we were happily returning a confidence for a partial window. It now reports warming_up instead.

None of these are model bugs. All four would have looked like model bugs if we had found them a week later, with real weights in the loop.

The cost, and what happens next

Two evenings of writing code that did nothing, and a habit of trusting the loop enough to believe the numbers coming out of it, once there are real numbers to believe. The plan for weeks five and six is to swap the stubs out one at a time: Aaron's CNN replaces features(), then the LSTM replaces confidence(). Nay Chi has already written up the interface contract the two of them will need to agree on before that happens, so we are not starting that argument from nothing.

I would do this again, and earlier. Writing the loop first forced us to decide what a frame, a window and an alert actually are while the decisions were still cheap, before there was a trained model around to make us feel like the hard part was already done.