Four silent failures a demo on somebody else's Mac walked straight into

Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.

## behavision-setup chose the Python least likely to work

findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".

maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.

## numpy<2.0 was the cap; OpenCV was the hazard

Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:

  numpy 1.26 / cv2 4.11    9 passed, 0 crashed
  numpy 2.0  / cv2 4.11    8 passed, 1 crashed
  numpy 1.26 / cv2 4.14    3 passed, 6 crashed
  numpy 2.0  / cv2 4.14    2 passed, 7 crashed

numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.

It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.

opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.

## One MQTT client id for a whole shop, so two PCs fought over it

behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:

  broker connected / broker connection lost: EOF / broker connected / EOF ...

The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.

MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.

## "no such file or directory" for an engine nobody had installed

Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.

engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-30 17:09:25 +05:30
parent 48a30d97db
commit ff4f95c3b0
11 changed files with 577 additions and 22 deletions

View File

@@ -1,7 +1,32 @@
"""cv2.FaceDetectorYN caches its input size and is not thread-safe, so camera
workers must not share one. Skipped when the model is absent, matching the
faiss-optional pattern in test_index.py — the suite stays runnable with no
models installed."""
models installed.
The premise is proved in a SUBPROCESS, and that is the whole lesson of this
file. The first version raced a shared detector in-process and asserted that
an exception came back, because on OpenCV 4.11 one usually did. It is a data
race in C++: what it produces is undefined, and on 4.14 what it mostly
produces is a segmentation fault. Measured, nine runs each, same machine:
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
numpy is not the variable; OpenCV is, and 4.11 crashing once says the hazard
was always there and 4.11 merely survived it. A test that takes the whole
suite down two runs in three is worse than no test: it turns "we upgraded
OpenCV" into a CI failure with no failing assertion in it, which is the
hardest kind to read.
None of this reaches the product. `Engine._build_worker` constructs a
FaceDetector per camera, which is what the rule says and what the second test
here guards.
"""
import subprocess
import sys
import textwrap
import threading
from pathlib import Path
@@ -11,10 +36,16 @@ import pytest
from behavision.config import Config
from behavision.detection import YUNET_FILENAME, FaceDetector
MODELS = Path(__file__).resolve().parent.parent / "models"
ROOT = Path(__file__).resolve().parent.parent
MODELS = ROOT / "models"
pytestmark = pytest.mark.skipif(not (MODELS / YUNET_FILENAME).exists(),
reason="YuNet model not installed")
# A shared detector fails probabilistically, so one clean attempt proves
# nothing. Several do: the premise holds if ANY attempt misbehaves, and only
# an unbroken run of clean ones is evidence it has stopped being true.
ATTEMPTS = 6
def _detector():
d = Config().detection
@@ -44,11 +75,37 @@ def _race(det_a, det_b):
return errors
_SHARED_RACE = textwrap.dedent("""
import sys
sys.path.insert(0, {root!r})
from tests.test_detector_concurrency import _detector, _race
shared = _detector()
print("raced" if _race(shared, shared) else "clean")
""")
def _shared_race_outcome():
"""Run one shared-detector race out of process.
Returns "raced" (an exception came back), "crashed" (the process died,
which is the same premise arriving by a blunter route) or "clean".
"""
proc = subprocess.run([sys.executable, "-c", _SHARED_RACE.format(root=str(ROOT))],
capture_output=True, text=True, timeout=120)
if proc.returncode != 0:
return "crashed"
return proc.stdout.strip().splitlines()[-1] if proc.stdout.strip() else "clean"
def test_a_shared_detector_really_does_race():
"""Guards the premise: if this ever stops failing, the test below is
proving nothing and the per-camera split can be revisited."""
shared = _detector()
assert _race(shared, shared), "expected a shared detector to race"
seen = [_shared_race_outcome() for _ in range(ATTEMPTS)]
assert any(o != "clean" for o in seen), (
f"a shared detector survived {ATTEMPTS} races ({seen}) - if that is "
"reproducible, cv2.FaceDetectorYN may have become thread-safe and the "
"per-camera rule in CLAUDE.md can be revisited"
)
def test_per_camera_detectors_do_not_race():