Four silent failures a demo on somebody else's Mac walked straight into

Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.

## behavision-setup chose the Python least likely to work

findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".

maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.

## numpy<2.0 was the cap; OpenCV was the hazard

Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:

  numpy 1.26 / cv2 4.11    9 passed, 0 crashed
  numpy 2.0  / cv2 4.11    8 passed, 1 crashed
  numpy 1.26 / cv2 4.14    3 passed, 6 crashed
  numpy 2.0  / cv2 4.14    2 passed, 7 crashed

numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.

It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.

opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.

## One MQTT client id for a whole shop, so two PCs fought over it

behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:

  broker connected / broker connection lost: EOF / broker connected / EOF ...

The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.

MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.

## "no such file or directory" for an engine nobody had installed

Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.

engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-30 17:09:25 +05:30
parent 48a30d97db
commit ff4f95c3b0
11 changed files with 577 additions and 22 deletions

107
CLAUDE.md
View File

@@ -3436,3 +3436,110 @@ three, and `api.CameraState` is the one function that decides them:
- **An unparseable `last_seen_at` is stale**, not connected. It should be
impossible, which is precisely why it must not fall through to the state that
says everything is fine.
## A demo on somebody else's Mac found four things, all of them silent
Three failures in one afternoon on a colleague's machine, plus one the fixing
uncovered. Every one produced a message that was true and useless.
### behavision-setup chose the Python least likely to work
`findPython` walked `3.14, 3.13, 3.12, 3.11, 3.10` and took the first hit — a
floor with **no ceiling**, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels for anything. It picked
3.14, pip found no numpy wheel for cp314 (`numpy<2.0` caps the resolver at
1.26.4, whose newest is cp312), fell back to building numpy from source and
produced `ERROR: Unknown compiler(s)`; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
`<arm_neon.h> is intended only for ARM and AArch64 targets`.
Two screens of C compiler output on a shop counter, for a version choice this
program made silently. `maxMinor` refuses in one line before anything is
downloaded, and **"too new" is a different message from "too old"** — telling
somebody holding Python 3.14 that no Python was found sends them to install a
newer one, which is the direction that just failed. It is a *wheel-availability*
ceiling, not a language one: onnxruntime is the binding dependency today
(cp314 is its newest), numpy publishes further ahead, and opencv ships a
stable-ABI wheel that covers everything.
### `numpy<2.0` was the cap; OpenCV was the hazard
Widening to `<3.0` needed proof, and the proof found something else. Nine runs
of the detector guard per combination, one machine, one sitting:
```
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
```
**numpy is not the variable; OpenCV is** — the third row is numpy 1.26. The
crash was `test_a_shared_detector_really_does_race`, which races a shared
`cv2.FaceDetectorYN` on purpose to prove the per-camera rule. That is undefined
behaviour in C++: 4.11 usually turned it into an exception, 4.14 usually turns
it into a **segfault**, and 4.11 crashing once says the hazard was always there
and 4.11 merely survived it.
It never reached the product — `Engine._build_worker` builds a detector per
camera, which is the rule and is what the second test guards. What it reached
was the suite: two runs in three died with **no failing assertion in them**,
turning "we upgraded OpenCV" into the hardest kind of CI failure to read. The
race now runs in a **subprocess**, so a segfault is an observed outcome rather
than the end of the run, and one clean attempt proves nothing — the premise
holds if *any* of several attempts misbehaves. With that fixed the suite is
226 passed / 2 skipped on numpy 2.0.2, five runs out of five.
`opencv-python` stays capped below 5. Everything above was measured on 4.x, and
an uncapped `>=4.8.1` means every NEW install silently gets a major release
this project has never run a real camera through while every existing one keeps
4.11.
### One MQTT client id for a whole shop, so two PCs fought over it
`behavision-<client>-<site>` is the same string on every computer claimed to
one site. MQTT requires client ids to be unique and a broker enforces it by
disconnecting the older session when a new one arrives with the same id, so the
colleague's Mac and the shop's own till took turns kicking each other off:
```
broker connected / broker connection lost: EOF / broker connected / EOF / ...
```
**The damage is not confined to the new machine.** The shop's till is the other
half of that loop, so somebody signing in on a laptop to look at the product
stops a live shop delivering visits — and from each end it reads as an unstable
network, because nothing says otherwise.
`Config.MQTTClientID()` appends a per-installation id, minted on first load and
written back so an existing install gets one without anybody doing anything.
The site stays in the name because that is what a broker log is read *by*. A
config that could not be written falls back to a per-run id rather than a
shared one: the right failure is a new name in the log after a restart, not the
collision this exists to end.
### "no such file or directory" for an engine nobody had installed
Pressing Start with no engine went straight to the supervisor, which reported
what `exec` reported:
```
engine failed to start: fork/exec /private/var/folders/c2/.../AppTranslocation/
500A5354-.../d/Behavision.app/Contents/MacOS/engine/behavision:
no such file or directory
```
Every word true, none of it saying *run the setup tool*. The startup path did
have that sentence — in a log file nobody on a shop counter opens.
`App.engineMissing()` is now the one function the startup path, the Start
button and the status panel all consult, so three surfaces cannot give three
accounts of one fact.
It also names **App Translocation**, which is in that path and is unguessable.
macOS quarantines a downloaded app it cannot verify and runs it from a randomly
named read-only copy, so every relative path resolves inside that copy — which
is why the engine folder appears missing from a bundle that plainly contains
one, and why installing into it would not survive a restart. Fixed by dragging
the app to Applications; saying nothing leaves somebody re-running a setup tool
that cannot win. The product is unsigned, so this is the *normal* first-run
state on every Mac, not an edge case.