7 Commits

Author SHA1 Message Date
b296e8a74a A demo does not need a tunnel; it needs the laptop's own camera
Asked after the mobile-internet question: could a VPN let the office cameras
be shown in the demo. Three jobs get confused there and only one needs one.

Seeing the estate from anywhere already works and needs nothing - viewer mode
plus LiveHub is exactly that, outbound, no installation and no credential.

Demonstrating recognition is better done on the demo machine's own camera.
`webcam: 0` picks a capture index instead of building an RTSP URL and the
engine has supported it since the first version: CameraStore round-trips it,
source() returns the index, safe_url() reports webcam:0, and
POST /api/cameras {"id":"laptop","webcam":0} has always worked. No screen
offered it - the same gap this repo already records for the customer record
and per-camera tuning. It is now an option in the make picker, and it is the
strongest demo available: real faces, in the room, depending on no network.
A demo pointed at a camera in another building depends on two internet
connections and a tunnel staying up while somebody is talking.

The address and the index are alternatives, not extras: source() takes the
webcam first, so a half-typed host left behind would make the saved camera
describe two things and use one. The scan, the path and the camera password
are hidden for a local camera because none of them mean anything.

A tunnel is still right for one case - running the engine on a remote machine
against the office's own cameras - and still wrong for the product: it is
per-site infrastructure on every shop PC, and it gives head office
network-level access into a customer's LAN, where today we can read a
camera's picture and nothing else.

Also: installing httpx took the suite from 239 passed to 271. The HTTP tests
importorskip it so a bare checkout runs, which means the number at the bottom
of a run is not the number of tests that exist. Added to the dev extra.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 18:18:57 +05:30
c9be9b5807 "The cameras won't connect from my mobile internet" was answered by a timeout
Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics and the product was not giving it.

A camera lives on the shop's LAN behind a router. 192.168.1.121 means
"something on the network I am attached to" and nothing more - from mobile
data, a hotel or head office it resolves to nobody, or to a completely
different device holding that number. There is no route in from the internet
and there must not be: an RTSP camera reachable from outside is how a shop's
cameras end up being watched by strangers.

That is why the product is split the way it is - the shop PC is the only
machine on the camera's LAN, and every other surface reaches it outbound,
which is what makes Watch live work from anywhere while nothing connects in.

What was wrong is the message. "cannot reach 192.168.1.121:554 - Operation
timed out" reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. _wrong_network_hint names the cause
and separates two states that need opposite actions:

  on that network   -> check the camera is powered on and the address is right
  somewhere else    -> the COMPUTER is in the wrong place; no setting fixes it

- The local address comes from a connected UDP socket that sends nothing. It
  only fixes a route so the kernel will name the source address.
- The LAN ranges are spelled out, not is_private. That property also covers
  carrier-grade NAT and the documentation networks, and telling somebody who
  typed 203.0.113.9 that it is "on the shop's own network" is a confident
  wrong answer in the place people look first. Found by a test using that
  address as its example of a PUBLIC one.
- A DNS name gets no hint: nothing can be concluded about camera.local from
  the string, and guessing is the failure mode this message exists to fix.
- With no network at all it still names the cause and drops the comparison,
  rather than claiming to know which network this machine is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 18:00:44 +05:30
0558344dc2 urllib does not let SSLCertVerificationError out, and a fake said it did
The certificate fix shipped and failed on the machine it was written for, with
the exact traceback it was meant to prevent. The retry was written

    except ssl.SSLCertVerificationError:

and urllib never raises that from urlopen. It catches it and re-raises
urllib.error.URLError(err), carrying the original on .reason. So the except
matched nothing, ever, and the fallback could not fire.

The unit test passed throughout, because the stub it used raised the bare SSL
error - a shape real urllib never produces. That is the lesson: a fake that
agrees with the author is worse than no test, because it converts an untested
path into a tested-looking one. This file already says that about
UPDATE ... RETURNING and about the in-memory API fake, and it got written
again anyway.

_is_cert_failure checks the exception and its .reason, and the tests now raise
URLError(SSLCertVerificationError(...)) - what the traceback actually shows. A
plain URLError is re-raised untouched, and a test asserts no second attempt is
made for one.

Beside the stubs there is now a real reproduction, opt-in behind
BEHAVISION_NETWORK_TESTS=1. Python's default context honours SSL_CERT_FILE, so
an empty file gives a context that trusts nobody - the python.org condition
exactly - while certifi is loaded by path and is unaffected. It skips rather
than passes where it cannot reproduce that, and the difference is measured:

    macOS Command Line Tools  LibreSSL 2.8.3   128 CAs with an empty CA file
    python.org / pyenv build  OpenSSL 3.5.8      0 CAs -> reproduces it

Checked for teeth by putting the shipped except back: both the corrected stub
test and the live one fail, and pass again when it is restored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:36:53 +05:30
8f07dee048 The engine stopped updating, and said "installed" every time
The SSL fix shipped and did not reach the machine it was written for. That
log said so, one line above the tick:

  behavision is already installed with the same version as the provided
  wheel. Use --force-reinstall to force an installation of the wheel.
   [ok] Engine and dependencies installed

and the traceback below it still pointed at model_assets.py line 59,
urllib.request.urlretrieve - code the fix had deleted.

Two frozen literals caused it: version = "1.1.0" in pyproject.toml and
__version__ = "1.0.0" in behavision/__init__.py. They disagreed with each
other and neither tracked a release, so every release built
behavision-1.1.0-py3-none-any.whl and pip install --upgrade on a machine that
already had 1.1.0 is a no-op. The comment beside that call claimed the
opposite.

The shape of the damage is what makes it bad. The Go binaries - app, agent,
setup tool - are rebuilt every release and updated normally. So a shop PC ran
a current app supervising an engine several releases old, and nothing said
which: /api/health reported the model, the paths, the cameras and the gallery,
and no version at all.

- One version, in the package, read by pyproject through
  [tool.setuptools.dynamic]. In a checkout it reads 0.0.0+dev: a
  plausible-looking number on a developer's /api/health is worse than none.
- pip install --force-reinstall --no-deps <wheel>, after the ordinary
  --upgrade. --upgrade settles the dependencies; the second call guarantees our
  own code is the code in the folder. --no-deps keeps it cheap - forcing the
  dependencies too would re-download ~300 MB every run. A rebuild at an
  unchanged version is the ordinary case while developing, so this must not
  rely on the version moving.
- release.sh stamps the tag: v0.5.6-demo -> 0.5.6+demo, valid PEP 440. Into a
  copy of the line, reverted in a trap, so the tree is never left dirty.

/api/health reports version now. Without it there is no way to tell a shop PC
three releases behind from a current one, which is how this survived several
releases.

Reproduced end to end before fixing, against real wheels on Python 3.12: two
builds of the same version, --upgrade leaves the old code in place and prints
the same sentence the colleague's Mac printed, --force-reinstall --no-deps
replaces it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:30:24 +05:30
3cddd9c2e1 The install finished and then could not download a 230 KB file
With the version ceiling and the widened numpy pin in place, setup succeeded
on the Mac that found them - Python 3.14 chosen and accepted, numpy 2.5.3,
onnxruntime 1.30, faiss 1.15.1, the engine itself - and died on the last step,
fetching the YuNet model:

  ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
  certificate verify failed: unable to get local issuer certificate

A python.org macOS build ships its own OpenSSL with NO trust store, and
populates one only when somebody double-clicks Install Certificates.command in
the Python folder. Nobody installing face-recognition software has a reason to
know that exists, and the failure is forty lines of traceback about _ssl.c at
the end of a ten-minute install.

_urlopen tries the default context first and retries with certifi's bundle on
a verification failure. The order is the design:

- Default first, because on Windows and on a system or Homebrew Python the
  default context reads the machine's own certificate store, which is what
  makes a corporate proxy with its own root CA work. Replacing it
  unconditionally would break every site that has one to fix a different
  platform.
- certifi second, because it is already installed: requests is a hard
  dependency and brings it.
- URLError is re-raised untouched. "No route to host" and "no trust store" are
  different problems, and retrying the first with a different CA list only
  delays the real message.

urlretrieve had to go, since it offers no way to pass a context - exactly the
kind of rewrite that silently drops something. The `download: <label> <n>%`
lines are a contract: supervisor.go's progressRe parses them to put first-run
progress in the tray, because the API is not up yet and a shop PC showing a
stopped engine for five minutes looks broken. A test asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:21:41 +05:30
248025cdf9 The Python ceiling was defeated by the venv the failure left behind
makeVenv reused any environment already on disk, whatever Python built it.
The machine that found the version bug already had a runtime built by 3.14,
left there by the run that failed - so with the ceiling in place setup would
choose a good interpreter, reach makeVenv, find the 3.14 environment, keep it,
and die in the same clang error as before.

A fix a user cannot reach because the bug's own debris is in the way is not a
fix, and it would have read as the release not working.

It now asks the interpreter inside an existing environment what it is and
rebuilds when the answer is unsupported, saying so. Rebuilding costs a
re-download of the libraries and nothing else - the models live in the state
root. An environment that cannot be asked counts as unusable too: a
half-created one answers nothing, and reusing it fails later in pip with an
error about a package rather than about the environment.

Tested against real environments rather than a fake, because what is under
test is what an interpreter on disk reports about itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:11:47 +05:30
ff4f95c3b0 Four silent failures a demo on somebody else's Mac walked straight into
Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.

## behavision-setup chose the Python least likely to work

findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".

maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.

## numpy<2.0 was the cap; OpenCV was the hazard

Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:

  numpy 1.26 / cv2 4.11    9 passed, 0 crashed
  numpy 2.0  / cv2 4.11    8 passed, 1 crashed
  numpy 1.26 / cv2 4.14    3 passed, 6 crashed
  numpy 2.0  / cv2 4.14    2 passed, 7 crashed

numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.

It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.

opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.

## One MQTT client id for a whole shop, so two PCs fought over it

behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:

  broker connected / broker connection lost: EOF / broker connected / EOF ...

The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.

MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.

## "no such file or directory" for an engine nobody had installed

Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.

engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:09:25 +05:30
23 changed files with 1554 additions and 102 deletions

366
CLAUDE.md
View File

@@ -3436,3 +3436,369 @@ three, and `api.CameraState` is the one function that decides them:
- **An unparseable `last_seen_at` is stale**, not connected. It should be
impossible, which is precisely why it must not fall through to the state that
says everything is fine.
## A demo on somebody else's Mac found four things, all of them silent
Three failures in one afternoon on a colleague's machine, plus one the fixing
uncovered. Every one produced a message that was true and useless.
### behavision-setup chose the Python least likely to work
`findPython` walked `3.14, 3.13, 3.12, 3.11, 3.10` and took the first hit — a
floor with **no ceiling**, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels for anything. It picked
3.14, pip found no numpy wheel for cp314 (`numpy<2.0` caps the resolver at
1.26.4, whose newest is cp312), fell back to building numpy from source and
produced `ERROR: Unknown compiler(s)`; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
`<arm_neon.h> is intended only for ARM and AArch64 targets`.
Two screens of C compiler output on a shop counter, for a version choice this
program made silently. `maxMinor` refuses in one line before anything is
downloaded, and **"too new" is a different message from "too old"** — telling
somebody holding Python 3.14 that no Python was found sends them to install a
newer one, which is the direction that just failed. It is a *wheel-availability*
ceiling, not a language one: onnxruntime is the binding dependency today
(cp314 is its newest), numpy publishes further ahead, and opencv ships a
stable-ABI wheel that covers everything.
### `numpy<2.0` was the cap; OpenCV was the hazard
Widening to `<3.0` needed proof, and the proof found something else. Nine runs
of the detector guard per combination, one machine, one sitting:
```
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
```
**numpy is not the variable; OpenCV is** — the third row is numpy 1.26. The
crash was `test_a_shared_detector_really_does_race`, which races a shared
`cv2.FaceDetectorYN` on purpose to prove the per-camera rule. That is undefined
behaviour in C++: 4.11 usually turned it into an exception, 4.14 usually turns
it into a **segfault**, and 4.11 crashing once says the hazard was always there
and 4.11 merely survived it.
It never reached the product — `Engine._build_worker` builds a detector per
camera, which is the rule and is what the second test guards. What it reached
was the suite: two runs in three died with **no failing assertion in them**,
turning "we upgraded OpenCV" into the hardest kind of CI failure to read. The
race now runs in a **subprocess**, so a segfault is an observed outcome rather
than the end of the run, and one clean attempt proves nothing — the premise
holds if *any* of several attempts misbehaves. With that fixed the suite is
226 passed / 2 skipped on numpy 2.0.2, five runs out of five.
`opencv-python` stays capped below 5. Everything above was measured on 4.x, and
an uncapped `>=4.8.1` means every NEW install silently gets a major release
this project has never run a real camera through while every existing one keeps
4.11.
### And the fix was defeated by the wreckage of the bug
`makeVenv` reused any environment already on disk, whatever Python built it.
That machine had a runtime built by **3.14**, left behind by the run that
failed — so with the ceiling in place setup would choose a good interpreter,
reach `makeVenv`, find the 3.14 environment, keep it, and die in the same clang
error as before. A fix a user cannot reach because the bug's own debris is in
the way is not a fix, and it would have read as the release not working.
It now asks the interpreter inside an existing environment what it is and
rebuilds when the answer is unsupported, saying so. Rebuilding costs a
re-download of the libraries and nothing else — the models live in the state
root, not in there. An environment that cannot be asked counts as unusable
too: a half-created one answers nothing, and reusing it fails later in pip
with an error about a package rather than about the environment.
### One MQTT client id for a whole shop, so two PCs fought over it
`behavision-<client>-<site>` is the same string on every computer claimed to
one site. MQTT requires client ids to be unique and a broker enforces it by
disconnecting the older session when a new one arrives with the same id, so the
colleague's Mac and the shop's own till took turns kicking each other off:
```
broker connected / broker connection lost: EOF / broker connected / EOF / ...
```
**The damage is not confined to the new machine.** The shop's till is the other
half of that loop, so somebody signing in on a laptop to look at the product
stops a live shop delivering visits — and from each end it reads as an unstable
network, because nothing says otherwise.
`Config.MQTTClientID()` appends a per-installation id, minted on first load and
written back so an existing install gets one without anybody doing anything.
The site stays in the name because that is what a broker log is read *by*. A
config that could not be written falls back to a per-run id rather than a
shared one: the right failure is a new name in the log after a restart, not the
collision this exists to end.
### "no such file or directory" for an engine nobody had installed
Pressing Start with no engine went straight to the supervisor, which reported
what `exec` reported:
```
engine failed to start: fork/exec /private/var/folders/c2/.../AppTranslocation/
500A5354-.../d/Behavision.app/Contents/MacOS/engine/behavision:
no such file or directory
```
Every word true, none of it saying *run the setup tool*. The startup path did
have that sentence — in a log file nobody on a shop counter opens.
`App.engineMissing()` is now the one function the startup path, the Start
button and the status panel all consult, so three surfaces cannot give three
accounts of one fact.
It also names **App Translocation**, which is in that path and is unguessable.
macOS quarantines a downloaded app it cannot verify and runs it from a randomly
named read-only copy, so every relative path resolves inside that copy — which
is why the engine folder appears missing from a bundle that plainly contains
one, and why installing into it would not survive a restart. Fixed by dragging
the app to Applications; saying nothing leaves somebody re-running a setup tool
that cannot win. The product is unsigned, so this is the *normal* first-run
state on every Mac, not an edge case.
### And then it could not download a 230 KB file
With all of the above fixed the install succeeded on that Mac - Python 3.14
chosen and accepted, numpy 2.5.3, onnxruntime 1.30, faiss 1.15.1, the engine
itself - and setup died on the last step, fetching the YuNet model:
```
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
certificate verify failed: unable to get local issuer certificate
```
A python.org macOS build ships its **own OpenSSL with no trust store**, and
populates one only when somebody double-clicks `Install Certificates.command`
in the Python folder. Nobody installing face-recognition software has any
reason to know that exists, and the failure is forty lines of traceback about
`_ssl.c` at the end of a ten-minute install.
`_urlopen` tries the default context first and retries with **certifi's**
bundle on a verification failure. The order is the whole design:
- Default first, because on Windows and on a system or Homebrew Python the
default context reads the machine's own certificate store - which is what
makes a corporate proxy with its own root CA work. Replacing it
unconditionally would break every site that has one in order to fix a
different platform.
- certifi second, because it is already installed: `requests` is a hard
dependency and brings it.
- A `URLError` is re-raised untouched. "No route to host" and "no trust store"
are different problems, and retrying the first with a different CA list only
delays the real message.
`urlretrieve` had to go, since it offers no way to pass a context - and that is
exactly the kind of rewrite that silently drops something. The
`download: <label> <n>%` lines are a **contract**: `supervisor.go`'s
`progressRe` parses them to put first-run progress in the tray, because the API
is not up yet and a shop PC showing a stopped engine for five minutes after
install looks broken. `tests/test_model_download.py` asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.
## The engine stopped updating, and said "installed" every time
The SSL fix above was released and **did not reach the machine it was written
for**. Its log said so, one line above the tick:
```
behavision is already installed with the same version as the provided wheel.
Use --force-reinstall to force an installation of the wheel.
[ok] Engine and dependencies installed
```
and the traceback that followed still pointed at `model_assets.py", line 59,
in _fetch / urllib.request.urlretrieve` — code the fix had deleted.
Two frozen literals caused it: `version = "1.1.0"` in `pyproject.toml` and
`__version__ = "1.0.0"` in `behavision/__init__.py`. They disagreed with each
other and neither tracked a release, so **every release built
`behavision-1.1.0-py3-none-any.whl`** and `pip install --upgrade` on a machine
that already had 1.1.0 is a no-op. The comment beside that call even claimed
the opposite: *"`--upgrade` so re-running after a new release replaces the
engine rather than leaving the old one in place and reporting success."*
The shape of the damage is what makes it bad. The Go binaries — the app, the
agent, the setup tool — are rebuilt every release and updated normally. So a
shop PC ran a current app supervising an engine several releases old, and
nothing anywhere said which: `/api/health` reported the model, the paths, the
cameras and the gallery, and **no version at all**.
Three changes, and the second is the one that does not depend on the first
being remembered:
- **One version, in the package**, read by `pyproject.toml` through
`[tool.setuptools.dynamic]`. Two literals in two files is how they came to
disagree. In a checkout it reads `0.0.0+dev`, because a plausible-looking
number on a developer's `/api/health` would be worse than none.
- **`pip install --force-reinstall --no-deps <wheel>`, after the ordinary
`--upgrade`.** `--upgrade` settles the dependencies; the second call
guarantees our own code is the code in the folder. `--no-deps` is what keeps
it cheap — forcing the dependencies too would re-download ~300 MB of numpy,
OpenCV and onnxruntime on every run. A rebuild at an unchanged version is the
ordinary case while developing, so this must not rely on the version moving.
- **`release.sh` stamps the tag**: `v0.5.6-demo` → `0.5.6+demo`, valid PEP 440
(a local segment takes alphanumerics and dots, never hyphens). Into a copy of
the line, reverted in a `trap`, so the tree is never left dirty.
`/api/health` reports `version` now. Without it there is no way to tell a shop
PC three releases behind from a current one, which is precisely how this
survived several releases.
Reproduced end to end before fixing, against real wheels on Python 3.12: two
builds of the same version, `--upgrade` leaves the old code in place and prints
the same sentence the colleague's Mac printed, `--force-reinstall --no-deps`
replaces it.
### `urllib` does not let `SSLCertVerificationError` out, and a fake said it did
The certificate fix above shipped and **failed on the machine it was written
for, with the exact traceback it was meant to prevent**. The retry was written
```python
except ssl.SSLCertVerificationError:
```
and urllib never raises that from `urlopen`. It catches it and re-raises
`urllib.error.URLError(err)`, carrying the original on `.reason`. So the
`except` matched nothing, ever, and the fallback could not fire.
The unit test passed throughout, because the stub it used raised the bare SSL
error — **a shape real urllib never produces**. That is the whole lesson: a
fake that agrees with the author is worse than no test, because it converts an
untested path into a tested-looking one. The same sentence is already in this
file about `UPDATE ... RETURNING` and about the in-memory API fake, and it was
written again here anyway.
`_is_cert_failure` checks the exception **and** its `.reason`, and the tests
now raise `URLError(SSLCertVerificationError(...))`, which is what his
traceback shows. A plain `URLError` — "no route to host" — is re-raised
untouched: retrying that with a different CA list changes nothing except how
long the operator waits for the real message, and a test asserts the second
attempt never happens.
Beside the stubs there is now a **real** reproduction,
`test_against_a_real_machine_with_no_trust_store`, opt-in behind
`BEHAVISION_NETWORK_TESTS=1` because it reaches github.com. Python's default
context honours `SSL_CERT_FILE`, so pointing it at an empty file gives a
default context that trusts nobody — the python.org condition exactly — while
certifi is loaded by path and is unaffected.
It skips rather than passes where it cannot reproduce that, and the difference
is measured rather than assumed:
```
macOS Command Line Tools LibreSSL 2.8.3 128 CAs with SSL_CERT_FILE empty
(reads the system keychain)
python.org / pyenv build OpenSSL 3.5.8 0 CAs -> reproduces it
```
Checked for teeth by putting the shipped `except` back: both the corrected stub
test and the live one fail, and pass again when it is restored.
## "The cameras won't connect from my mobile internet" — and why that is not a bug
Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics, and the product was not giving it.
A camera lives on the shop's LAN behind a router. `192.168.1.121` means
*"something on the network I am attached to"* and nothing more — from mobile
data, a hotel, or head office it resolves to nobody, or to a completely
different device that happens to hold that number. There is no route in from
the internet and **there must not be**: an RTSP camera reachable from outside
is how a shop's cameras end up being watched by strangers.
This is the reason the product is split the way it is, and it is worth stating
plainly next to the split itself: the shop PC is the only machine on the
camera's LAN, so it does the connecting, and every other surface reaches it
**outbound** — the agent's pull, the arrivals feed, and `LiveHub`'s frame relay,
which is what makes **Watch live** work from anywhere while nothing connects in.
What was wrong is the message. `cannot reach 192.168.1.121:554 - Operation
timed out` reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. `_wrong_network_hint` now names the
cause, and it distinguishes two states that need opposite actions — the same
rule as `artifact` against `no_faces`, and `stale` against `not_connecting`:
| | |
|---|---|
| this computer **is** on that network | check the camera is powered on and that the address is right |
| this computer is **somewhere else** | the computer is in the wrong place; no setting here fixes it, recognition has to run on a machine in the shop |
- **The local address comes from a `connect`ed UDP socket** that sends nothing.
It only fixes a route so the kernel will name the source address — no packet
leaves, and it needs no dependency in an engine that already ships 200 MB of
models.
- **The LAN ranges are spelled out, not `is_private`.** That property is
broader than "an address on somebody's LAN": it also covers carrier-grade NAT
and the documentation networks (192.0.2, 198.51.100, 203.0.113), and telling
somebody who typed one of those that it is "on the shop's own network" is a
confident wrong answer in the place people look first. Found by a test using
`203.0.113.9` as its example of a *public* address, which `is_private` calls
private.
- **A DNS name gets no hint at all.** Nothing can be concluded about
`camera.local` from the string, and guessing is the failure mode this whole
message exists to fix.
- **With no network at all it still names the cause** and drops the comparison,
rather than claiming to know which network this machine is on.
## A tunnel for the demo: when it is the right answer, and when it is not
Asked after the mobile-internet question: *"what if we created a secure tunnel
or vpn, then the cameras in our office can be shown in the demo version too?"*
Three different jobs get confused here, and only one of them needs a tunnel.
**Seeing the estate from anywhere already works and needs nothing.** Viewer
mode plus `LiveHub` is exactly this: the shop PC pushes frames outbound and any
signed-in app sees them, from mobile data, a hotel or a customer's office. A
VPN would add an installation, a credential and a moving part to something that
already works with none of them.
**Demonstrating recognition is better done on the demo machine's own camera.**
`webcam: 0` picks a capture index instead of building an RTSP URL, and the
engine has supported it since the first version — `CameraStore` round-trips it,
`source()` returns the index, `safe_url()` reports `webcam:0`, and
`POST /api/cameras {"id":"laptop","webcam":0}` has always worked. **No screen
offered it**, which is the same gap this file already records for the customer
record and for per-camera tuning: the API could, the UI could not reach it.
It is now an option in the make picker, and it is the strongest demo available:
real faces, in the room, instant, depending on no network at all. A demo
pointed at a camera in another building depends on two internet connections and
a tunnel staying up while somebody is talking.
**The one case a tunnel genuinely answers** is running the engine on a remote
machine against the office's own cameras — LAN access to `192.168.1.121` from
somewhere that is not that LAN. A mesh VPN (Tailscale and similar) does that
honestly: a subnet router at the office, the client on the demo machine, and
the address works unchanged. Free at this scale, no port forwarding, no
exposed camera. Worth using for our *own* office when that is really the goal.
**It is still the wrong answer for the product**, and the reasons are not about
difficulty:
- It is per-site infrastructure — an account, a node, a key — on every shop PC
we ship, to replace something that already works over ordinary HTTPS.
- It gives head office **network-level access into a customer's LAN**. The
current design can read a camera's picture; a VPN can reach the customer's
till, their router and everything else on that network. That is a far larger
thing to be trusted with, and a far larger thing to have breached.
- The failure modes are worse and less legible: a tunnel that is down looks
like a camera that is down.
So: tunnel for our own office if we want the engine running remotely; nothing
at all for showing customers their shops; the laptop's own camera for showing
anybody what the product does.
### And the suite was quietly 32 tests smaller than it looked
Installing `httpx` to check the webcam path took the engine suite from **239
passed to 271**. `tests/test_api_cameras.py` and its siblings begin with
`pytest.importorskip("httpx")` so a bare checkout still runs — deliberate, and
it means the number at the bottom of a run is not the number of tests that
exist. `pip install -e .[dev]` is the opt-in.

View File

@@ -46,6 +46,61 @@ import (
// otherwise arrive as a syntax error deep inside a dependency.
const minMinor = 10
// maxMinor is a WHEEL-availability ceiling, not a language one, and it is the
// reason this constant exists at all.
//
// findPython used to take the newest interpreter it could find, with a floor
// and no ceiling - which is precisely backwards, because the newest Python is
// the one least likely to have binary wheels for anything. Measured on a
// second Mac: it chose Python 3.14, pip found no numpy wheel for cp314, fell
// back to building numpy from source, and produced
//
// ERROR: Unknown compiler(s): [['cc'], ['gcc'], ['clang'], ...]
//
// then, once the operator installed Xcode's command line tools to get past
// that, ten minutes of compiling ending in
//
// arm_neon.h:28:2: error: "<arm_neon.h> is intended only for ARM and
// AArch64 targets"
//
// Two screens of C compiler output, on a shop counter, for a version choice
// made silently by this program. Refusing in one line, before anything is
// downloaded, is the whole of the fix.
//
// Raise it when the dependency set has wheels for the next version. Today
// onnxruntime is the binding one (cp314 is its newest); numpy publishes
// further ahead, and opencv-python ships a stable-ABI wheel that covers
// everything. `pip download --only-binary=:all: -r requirements.txt` against
// a candidate interpreter is the check.
const maxMinor = 14
// The three answers a candidate interpreter can get. Three, not two: a
// version that is too new and one that is too old need opposite actions from
// the operator, and collapsing them tells somebody holding Python 3.14 to go
// and install a newer Python.
const (
verdictOK = "ok"
verdictTooOld = "old"
verdictTooNew = "new"
verdictUnknown = "unparseable"
)
func pythonVerdict(major, minor int, parsed bool) string {
switch {
case !parsed:
return verdictUnknown
case major != 3:
// Python 4 is not a version this has been tried against, and 2 is
// long gone. Neither is a thing to guess about.
return verdictTooNew
case minor < minMinor:
return verdictTooOld
case minor > maxMinor:
return verdictTooNew
}
return verdictOK
}
func main() {
if err := run(); err != nil {
fmt.Fprintf(os.Stderr, "\n Setup did not finish: %v\n\n", err)
@@ -267,7 +322,13 @@ func findPython() (string, string, error) {
// `python3` therefore told a Mac with Python 3.12 sitting on it to go and
// install Python - measured on this machine, which has 3.12 under
// ~/.local/opt and reported "Found, but too old: python3 3.9".
versions := []string{"3.14", "3.13", "3.12", "3.11", "3.10"}
// Newest first WITHIN the supported range. Newest overall is what broke
// this; a version nobody has built wheels for is not a better choice than
// one that works.
var versions []string
for v := maxMinor; v >= minMinor; v-- {
versions = append(versions, fmt.Sprintf("3.%d", v))
}
for _, v := range versions {
cands = append(cands, cand{"python" + v, nil})
}
@@ -290,7 +351,7 @@ func findPython() (string, string, error) {
}
}
var tried []string
var tried, tooNew []string
for _, c := range cands {
exe := c.exe
if filepath.IsAbs(exe) {
@@ -314,7 +375,18 @@ func findPython() (string, string, error) {
}
ver := strings.TrimSpace(string(out))
tried = append(tried, c.exe+" "+ver)
if major, minor, ok := parseVer(ver); ok && (major > 3 || (major == 3 && minor >= minMinor)) {
major, minor, parsed := parseVer(ver)
switch verdict := pythonVerdict(major, minor, parsed); verdict {
case verdictTooNew:
// Recorded separately: "too new" and "too old" need opposite
// actions, and a single "found, but unsuitable" list sends
// somebody to upgrade a Python that is already past the problem.
tooNew = append(tooNew, c.exe+" "+ver)
continue
case verdictTooOld, verdictUnknown:
continue
}
{
full := exe
if len(c.args) > 0 {
full = exe + " " + strings.Join(c.args, " ")
@@ -327,7 +399,30 @@ func findPython() (string, string, error) {
// python.exe to PATH" on a Windows installer page reads as software that
// does not know where it is running, which is exactly the moment somebody
// stops trusting the rest of what it says.
msg := "no Python 3.10 or newer was found on this computer.\n\n"
// Only a too-new Python is a different problem with a different fix, and
// saying "no Python was found" to somebody looking at Python 3.14 is the
// kind of message that makes people stop believing the next one.
if len(tooNew) > 0 && len(tried) == 0 {
// Built as a value and wrapped, not written as an fmt.Errorf literal:
// this is a paragraph shown to an operator, and a linter that wants
// error strings to be lower-case fragments is right about errors
// programs read and wrong about the ones people do.
tooNewMsg := fmt.Sprintf(
"this computer has %s, which is newer than Behavision supports.\n\n"+
" Some of the libraries the engine needs have no build for it\n"+
" yet, so installing would fail part-way through.\n\n"+
" Install Python 3.%d and run this again:\n"+
" macOS: brew install python@3.%d\n"+
" or https://www.python.org/downloads/macos/\n"+
" Windows: https://www.python.org/downloads/windows/\n\n"+
" Both versions can sit on the machine together; this picks\n"+
" the one it can use.",
strings.Join(tooNew, ", "), maxMinor, maxMinor)
return "", "", errors.New(tooNewMsg)
}
msg := fmt.Sprintf("no Python between 3.%d and 3.%d was found on this computer.\n\n",
minMinor, maxMinor)
if runtime.GOOS == "windows" {
msg += " Install it from https://www.python.org/downloads/windows/\n" +
" and tick \"Add python.exe to PATH\" on the first screen,\n" +
@@ -339,6 +434,9 @@ func findPython() (string, string, error) {
if len(tried) > 0 {
msg += "\n\n Found, but too old: " + strings.Join(tried, ", ")
}
if len(tooNew) > 0 {
msg += "\n\n Found, but too new: " + strings.Join(tooNew, ", ")
}
return "", "", errors.New(msg)
}
@@ -375,14 +473,52 @@ func venvPython(venv string) string {
// engine requires - inside a shared interpreter is how you break the other
// thing months later, silently.
func makeVenv(py, venv string) error {
// An existing environment is reused - but only if the Python inside it is
// one this build supports.
//
// It used to be reused unconditionally, and that would have made the
// version ceiling above look like it did not work. The machine this was
// all found on already had a runtime built by Python 3.14, from the run
// that failed: with the ceiling in place setup would choose a good
// interpreter, reach here, find the 3.14 environment, keep it, and die in
// the same clang error as before. A fix that is defeated by the wreckage
// of the bug it fixes is not one.
//
// Rebuilding costs a re-download of the libraries and nothing else. The
// models are in the state root, not in here, so they survive.
if _, err := os.Stat(venvPython(venv)); err == nil {
return nil // already built; pip below brings it up to date
ok, ver := venvUsable(venv)
if ok {
return nil // pip below brings it up to date
}
fmt.Printf(" [..] %-24s %s\n", "Rebuilding environment",
"the existing one uses "+ver+", which is not supported")
if err := os.RemoveAll(venv); err != nil {
return fmt.Errorf("removing the old environment at %s: %w", venv, err)
}
}
exe, args := splitLauncher(py)
args = append(args, "-m", "venv", venv)
return stream(exec.Command(exe, args...), "creating the virtual environment")
}
// venvUsable reports whether the interpreter already inside an environment is
// one this build supports, and what it is when it is not.
//
// An environment that cannot be asked counts as unusable: a half-created or
// truncated one answers nothing, and reusing it fails later in pip with an
// error about a package rather than about the environment.
func venvUsable(venv string) (bool, string) {
out, err := exec.Command(venvPython(venv), "-c",
"import sys;print('%d.%d'%sys.version_info[:2])").Output()
if err != nil {
return false, "an interpreter that will not run"
}
ver := strings.TrimSpace(string(out))
major, minor, parsed := parseVer(ver)
return pythonVerdict(major, minor, parsed) == verdictOK, "Python " + ver
}
func pipInstall(vpy, src string) error {
fmt.Println(" Installing the engine and its libraries. This downloads a few")
fmt.Println(" hundred megabytes and takes a while on a slow connection.")
@@ -403,7 +539,30 @@ func pipInstall(vpy, src string) error {
// 'behavision.egg-info': Read-only file system". Falling back to source
// copies it somewhere writable first, for the same reason.
if wheels, _ := filepath.Glob(filepath.Join(src, "behavision-*.whl")); len(wheels) > 0 {
return stream(exec.Command(vpy, "-m", "pip", "install", "--upgrade", wheels[0]),
// TWO calls, and the second is the one that matters.
//
// `--upgrade` alone is not an upgrade when the version has not moved:
// pip skips the wheel and says so, one line above this program
// printing "[ok] Engine and dependencies installed". The wheel version
// was a frozen literal for several releases, so every engine fix in
// them silently failed to reach any machine that had run setup once -
// while the Go binaries beside it, rebuilt every release, updated
// normally. Half the product current, half of it months old, and
// nothing saying which.
//
// release.sh stamps the tag into the version now, so the versions do
// differ. This does not rely on that: a rebuild at the same version is
// the ordinary case while developing, and "installed" has to mean the
// code in this folder either way.
if err := stream(exec.Command(vpy, "-m", "pip", "install", "--upgrade", wheels[0]),
"installing the engine"); err != nil {
return err
}
// --no-deps so this is our own package only: the call above has
// already settled the dependencies, and forcing those too would
// re-download ~300 MB of numpy, OpenCV and onnxruntime every run.
return stream(exec.Command(vpy, "-m", "pip", "install",
"--force-reinstall", "--no-deps", wheels[0]),
"installing the engine")
}
tmp, err := os.MkdirTemp("", "behavision-src-")

View File

@@ -0,0 +1,112 @@
package main
import (
"os"
"path/filepath"
"testing"
)
// The choice this program makes silently, and got wrong.
//
// findPython took the newest interpreter on the machine, with a floor and no
// ceiling - backwards, because the newest Python is the one least likely to
// have binary wheels. On a Mac holding Python 3.14 it chose 3.14, pip found
// no numpy wheel for cp314, fell back to a source build and produced two
// screens of clang errors ending in "<arm_neon.h> is intended only for ARM
// and AArch64 targets". The operator's machine was fine; the version was not.
func TestTooNewIsRefusedRatherThanCompiled(t *testing.T) {
if got := pythonVerdict(3, maxMinor+1, true); got != verdictTooNew {
t.Errorf("3.%d = %q, want %q - picking it means a source build",
maxMinor+1, got, verdictTooNew)
}
if got := pythonVerdict(3, maxMinor, true); got != verdictOK {
t.Errorf("3.%d = %q, want %q - the ceiling is inclusive", maxMinor, got, verdictOK)
}
}
// Too old and too new must stay different answers. Telling somebody holding
// Python 3.14 that no Python was found, or that theirs is too old, sends them
// to install a newer one - which is the direction that already failed.
func TestOldAndNewAreDifferentAnswers(t *testing.T) {
old := pythonVerdict(3, minMinor-1, true)
fresh := pythonVerdict(3, maxMinor+1, true)
if old == fresh {
t.Fatalf("3.%d and 3.%d both reported %q", minMinor-1, maxMinor+1, old)
}
if old != verdictTooOld {
t.Errorf("3.%d = %q, want %q", minMinor-1, old, verdictTooOld)
}
}
// Every version in the range is accepted, so the window this program claims
// to support is the one it actually uses.
func TestTheWholeSupportedRangeIsAccepted(t *testing.T) {
for m := minMinor; m <= maxMinor; m++ {
if got := pythonVerdict(3, m, true); got != verdictOK {
t.Errorf("3.%d = %q, want %q", m, got, verdictOK)
}
}
if minMinor > maxMinor {
t.Fatal("the supported range is empty; nothing would ever be chosen")
}
}
// A major version nobody has tested against is not something to guess at, and
// an unreadable version string is not a working interpreter.
func TestUnknownVersionsAreNotAccepted(t *testing.T) {
for _, c := range []struct {
name string
major, minor int
parsed bool
}{
{"python 4", 4, 0, true},
{"python 2", 2, 7, true},
{"unparseable", 0, 0, false},
} {
if got := pythonVerdict(c.major, c.minor, c.parsed); got == verdictOK {
t.Errorf("%s was accepted", c.name)
}
}
}
// An environment already on disk is reused, and that is right until the Python
// inside it is one this build cannot use.
//
// It was reused unconditionally, which would have defeated the ceiling above
// on the exact machine that found the bug: that Mac already had a runtime
// built by Python 3.14, left behind by the run that failed. Setup would pick a
// good interpreter, find the 3.14 environment, keep it, and die in the same
// clang error as before - a fix defeated by the wreckage of the bug it fixes.
//
// Real environments, not a fake: the thing under test is what an interpreter
// on disk reports about itself.
func TestAnUnsupportedEnvironmentIsNotReused(t *testing.T) {
py, _, err := findPython()
if err != nil {
t.Skipf("no supported Python on this machine: %v", err)
}
venv := filepath.Join(t.TempDir(), "runtime")
if err := makeVenv(py, venv); err != nil {
t.Fatalf("makeVenv: %v", err)
}
if ok, ver := venvUsable(venv); !ok {
t.Fatalf("an environment built from the interpreter setup just chose "+
"reported itself unusable (%s)", ver)
}
// The two states that must not be confused with a working one.
empty := filepath.Join(t.TempDir(), "gone")
if ok, _ := venvUsable(empty); ok {
t.Error("a missing environment was reported usable")
}
broken := filepath.Join(t.TempDir(), "broken")
if err := os.MkdirAll(filepath.Dir(venvPython(broken)), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(venvPython(broken), []byte("not an interpreter"), 0o755); err != nil {
t.Fatal(err)
}
if ok, ver := venvUsable(broken); ok {
t.Errorf("a half-created environment was reported usable (%s)", ver)
}
}

View File

@@ -325,7 +325,7 @@ func cmdRun() error {
cfg.ClientID, cfg.SiteID, cfg.BrokerURL)
client, err := mqtt.NewClient(mqtt.ClientOptions{
BrokerURL: cfg.BrokerURL,
ClientID: "behavision-" + cfg.ClientID + "-" + cfg.SiteID,
ClientID: cfg.MQTTClientID(),
Username: cfg.BrokerUsername, Password: cfg.BrokerPassword,
CAFile: cfg.BrokerCAFile, Log: logger,
})

View File

@@ -30,12 +30,12 @@ import (
// argument, and it is why the wanted-check comes first and the push stops the
// moment the server says the last viewer has gone.
type Live struct {
Engine *EngineClient
Cloud *CloudClient
Log *log.Logger
FPS float64
Width int
Quality int
Engine *EngineClient
Cloud *CloudClient
Log *log.Logger
FPS float64
Width int
Quality int
}
// Defaults, measured against the office camera rather than guessed.

View File

@@ -8,7 +8,9 @@
package config
import (
"crypto/rand"
"encoding/base64"
"encoding/hex"
"encoding/json"
"fmt"
"os"
@@ -83,6 +85,10 @@ type Config struct {
// Queue.
SpoolMax int `json:"spool_max"`
// InstallID distinguishes THIS installation from every other one claimed
// to the same site. See MQTTClientID.
InstallID string `json:"install_id,omitempty"`
path string
}
@@ -124,6 +130,14 @@ func Load(path string) (Config, error) {
return cfg, fmt.Errorf("config %s: %w", path, err)
}
cfg.path = path
// Minted on first load and written back, so an installation that predates
// this field gets one without anybody doing anything. Best effort: a
// read-only config still yields a working id for this run, it is simply
// not the same one next time.
if cfg.InstallID == "" {
cfg.InstallID = newInstallID()
_ = cfg.Save(path)
}
for _, field := range []*string{&cfg.BrokerPassword, &cfg.APIPassword,
&cfg.SessionToken, &cfg.SessionRefresh, &cfg.AgentToken} {
plain, err := reveal(*field)
@@ -210,3 +224,41 @@ func reveal(stored string) (string, error) {
}
return string(plain), nil
}
// MQTTClientID names this INSTALLATION, not this site.
//
// It was `behavision-<client>-<site>`, which is the same string on every
// computer claimed to one shop. MQTT requires client ids to be unique and a
// broker enforces it by disconnecting the older session when a new one
// arrives with the same id - so two machines on one site take turns kicking
// each other off, forever. Measured on a second Mac claimed to a live shop:
//
// broker connected / broker connection lost: EOF / broker connected / ...
//
// The damage is not confined to the new machine. The shop's own till is the
// other half of that loop, so somebody signing in on a laptop to look at the
// product stops the shop delivering visits - and nothing at either end says
// why, because from each side it reads as an unstable network.
//
// The site stays in the id because it is what a broker log is read by, and
// the random half is short for the same reason. `CleanSession(true)` means
// there is no session state for a changed id to strand.
func (c Config) MQTTClientID() string {
id := c.InstallID
if id == "" {
// A config that could not be written still has to produce a UNIQUE
// id, or this falls straight back into the collision it exists to
// prevent. Per-run is the right failure: the connection works and the
// only cost is a new name in the broker's log after a restart.
id = newInstallID()
}
return "behavision-" + c.ClientID + "-" + c.SiteID + "-" + id
}
func newInstallID() string {
b := make([]byte, 4)
if _, err := rand.Read(b); err != nil {
return "x"
}
return hex.EncodeToString(b)
}

View File

@@ -0,0 +1,88 @@
package config
import (
"path/filepath"
"strings"
"testing"
)
// The bug this exists to prevent, measured on a second Mac claimed to a live
// shop: MQTT requires client ids to be unique, and a broker enforces it by
// disconnecting the older session when a new one arrives with the same id. The
// id was `behavision-<client>-<site>` - identical on every computer claimed to
// one shop - so the two took turns kicking each other off:
//
// broker connected / broker connection lost: EOF / broker connected / ...
//
// The damage is not confined to the new machine. The shop's own till is the
// other half of that loop, so somebody signing in on a laptop to look at the
// product stops the shop delivering visits.
func TestTwoInstallsOnOneSiteGetDifferentClientIDs(t *testing.T) {
dir := t.TempDir()
one := writeClaimed(t, filepath.Join(dir, "a.json"))
two := writeClaimed(t, filepath.Join(dir, "b.json"))
if one.MQTTClientID() == two.MQTTClientID() {
t.Fatalf("both installs answered to %q; the broker will disconnect one "+
"whenever the other connects", one.MQTTClientID())
}
}
// And the same install keeps its name across restarts, or a broker log is a
// list of strangers and nobody can tell one till from a stream of new ones.
func TestOneInstallKeepsItsClientIDAcrossRestarts(t *testing.T) {
path := filepath.Join(t.TempDir(), "agent.json")
first := writeClaimed(t, path)
again, err := Load(path)
if err != nil {
t.Fatalf("reload: %v", err)
}
if got, want := again.MQTTClientID(), first.MQTTClientID(); got != want {
t.Errorf("after a restart the id was %q, want %q", got, want)
}
}
// The site stays in the id: it is what somebody reading a broker log is
// reading FOR, and an opaque random string would make every connection
// anonymous.
func TestTheClientIDStillNamesTheShop(t *testing.T) {
c := Config{ClientID: "tenext-retail", SiteID: "chennai", InstallID: "abcd1234"}
id := c.MQTTClientID()
for _, want := range []string{"tenext-retail", "chennai", "abcd1234"} {
if !strings.Contains(id, want) {
t.Errorf("client id %q does not contain %q", id, want)
}
}
}
// A config that could not be written still has to produce a UNIQUE id, or a
// read-only install falls straight back into the collision. Per-run is the
// right failure: the connection works, and the only cost is a new name in the
// broker's log after a restart.
func TestAnUnsavedConfigStillGetsAUniqueID(t *testing.T) {
a := Config{ClientID: "c", SiteID: "s"}
b := Config{ClientID: "c", SiteID: "s"}
if a.MQTTClientID() == b.MQTTClientID() {
t.Fatal("two configs with no install id produced the same client id")
}
}
func writeClaimed(t *testing.T, path string) Config {
t.Helper()
cfg := Defaults()
cfg.ClientID, cfg.SiteID = "tenext-retail", "chennai"
if err := cfg.Save(path); err != nil {
t.Fatalf("save: %v", err)
}
// Loading is what mints the id, so an installation that predates the
// field gets one without anybody doing anything.
got, err := Load(path)
if err != nil {
t.Fatalf("load: %v", err)
}
if got.InstallID == "" {
t.Fatal("loading a config without an install id did not mint one")
}
return got
}

View File

@@ -4,4 +4,15 @@ Pipeline: capture -> detect (YuNet) -> track (IoU) -> align + encode
(ArcFace ONNX) -> match / auto-enroll (FAISS + SQLite) -> events + API.
"""
__version__ = "1.0.0"
# Stamped by release.sh from the git tag, into a copy of this line, so a
# release always produces a wheel nobody has installed before. It was a frozen
# "1.0.0" here and a frozen "1.1.0" in pyproject.toml - two literals that
# disagreed with each other and tracked nothing - which is how every engine fix
# for several releases silently failed to reach a machine that had run setup
# once. pip skips a wheel whose version is already installed and says so, one
# line above setup printing "[ok] Engine and dependencies installed".
#
# In a checkout it stays obviously a checkout: "0.0.0+dev" on /api/health is
# the truth about a developer's machine, and a plausible-looking number there
# would be worse than none.
__version__ = "0.0.0+dev"

View File

@@ -17,6 +17,7 @@ from fastapi.responses import HTMLResponse, Response, StreamingResponse
from fastapi.security import HTTPBasic, HTTPBasicCredentials
from pydantic import BaseModel, ValidationError
from . import __version__
from .config import ApiSection, CameraConfig, CameraTuning
from .commission import CommissionRun
from .events import Event
@@ -142,7 +143,7 @@ def _reencode(jpeg: bytes, width: int, quality: int) -> "bytes | None":
def create_app(engine: Engine) -> FastAPI:
app = FastAPI(title="Behavision", version="1.0.0",
app = FastAPI(title="Behavision", version=__version__,
dependencies=_auth_dependencies(engine.cfg.api))
def worker_or_404(camera_id: str):
@@ -172,6 +173,10 @@ def create_app(engine: Engine) -> FastAPI:
# are up, and the shop recognises nobody it already knows.
stranded = engine.gallery.health["stranded"]
return {"status": "ok" if engine.started_at else "starting",
# Which build is actually running. Without it there was no way
# to tell a shop PC three releases behind from a current one -
# which is precisely how a silent install failure survived.
"version": __version__,
"recognition_model": engine.encoder.model_name,
"gallery_unreadable_embeddings": stranded,
# "where is my database" must be answerable from the API: the

View File

@@ -68,6 +68,78 @@ os.environ.setdefault(
STALL_AFTER_S = 10.0
def _local_ipv4() -> str:
"""This machine's address on the interface holding the default route.
A UDP socket is `connect`ed and nothing is sent - it only fixes a route so
the kernel will name the source address. No packet leaves, and it needs no
dependency, which matters in an engine that already ships 200 MB of models.
"""
import socket
try:
with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as s:
s.settimeout(0.5)
s.connect(("8.8.8.8", 80))
return s.getsockname()[0]
except OSError:
return ""
def _wrong_network_hint(host: str) -> str:
"""Why a private camera address is unreachable, when that is the reason.
A camera lives on the shop's LAN behind a router, and 192.168.x.x means
"something on the network I am attached to" - nothing more. From mobile
data, a hotel, or head office it either resolves to nobody or to a
completely different device that happens to hold that number. There is no
route in from the internet and there must not be: an RTSP camera reachable
from outside is how a shop's cameras end up being watched by strangers.
Without this the answer was "cannot reach 192.168.1.121:554 - Operation
timed out", which reads as a broken camera and sends somebody to re-type
an address and a password that were always correct. Asked directly by the
owner, about his own cameras, from his phone's connection.
Two states, two different actions, so they must not share a sentence: on
the same network the camera or its address is the problem; on a different
one the COMPUTER is in the wrong place and no setting will fix it.
"""
import ipaddress
try:
addr = ipaddress.ip_address(host)
except ValueError:
return "" # a DNS name; nothing can be concluded from the string
# The RFC1918 blocks and link-local, spelled out rather than `is_private`.
# That property is broader than "an address on somebody's LAN": it also
# covers the carrier-grade NAT range and the documentation networks
# (192.0.2, 198.51.100, 203.0.113), and telling somebody who typed one of
# those that it is "on the shop's own network" would be a confident wrong
# answer in the place people look first. Found by a test using 203.0.113.9
# as an example of a PUBLIC address, which `is_private` calls private.
lan = any(addr in ipaddress.ip_network(n) for n in
("10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "169.254.0.0/16")
if addr.version == 4)
if not lan:
return ""
mine = _local_ipv4()
if not mine:
return (" - that is a private address, reachable only from inside "
"the network the camera is on")
try:
same = ipaddress.ip_network(f"{mine}/24", strict=False).supernet_of(
ipaddress.ip_network(f"{host}/24", strict=False))
except (ValueError, TypeError):
same = False
if same:
return (f" - this computer is on that network ({mine}), so check the "
f"camera is powered on and that {host} is its address")
return (f" - this computer is on {mine}, not the camera's network. A "
f"private address like {host} is only reachable from inside the "
f"shop's own network, never over the internet or mobile data, so "
f"recognition has to run on a computer in the shop")
def _tcp_reachable(source: "str | int", timeout: float
) -> "tuple[bool, str]":
"""Cheap pre-flight for an rtsp:// URL. Non-URL sources pass through."""
@@ -85,10 +157,11 @@ def _tcp_reachable(source: "str | int", timeout: float
return True, ""
except socket.timeout:
return False, (f"no response from {parsed.hostname}:{port} within "
f"{timeout:.0f}s - check the IP address and that the "
f"camera is on the same network")
f"{timeout:.0f}s{_wrong_network_hint(parsed.hostname)}")
except OSError as exc:
return False, f"cannot reach {parsed.hostname}:{port} - {exc.strerror or exc}"
return False, (f"cannot reach {parsed.hostname}:{port} - "
f"{exc.strerror or exc}"
f"{_wrong_network_hint(parsed.hostname)}")
def _fourcc(cap) -> str:

View File

@@ -4,6 +4,8 @@ from __future__ import annotations
import logging
import shutil
import ssl
import urllib.error
import urllib.request
from pathlib import Path
@@ -35,6 +37,81 @@ _COPY_MAP = {
}
def _https_context() -> "ssl.SSLContext | None":
"""The CA store to trust, or None to use whatever Python defaults to.
Returning None first is deliberate. On Windows and on a Homebrew or
system Python, the default context reads the machine's own certificate
store - which is what makes a corporate proxy with its own root CA work.
Replacing that with certifi's bundle unconditionally would break every
site that has one, in order to fix a different platform.
The platform this fixes is a python.org macOS build. It ships its own
OpenSSL with NO trust store, and populates one only when somebody
double-clicks `Install Certificates.command` in the Python folder -
which nobody installing face-recognition software has any reason to know
about. Every HTTPS request from that interpreter fails with:
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
certificate verify failed: unable to get local issuer certificate
Measured on a colleague's Mac: the engine installed perfectly and then
could not download a 230 KB model file, ending setup in forty lines of
traceback about `_ssl.c`.
"""
try:
import certifi
except ImportError: # pragma: no cover - certifi ships with requests
return None
return ssl.create_default_context(cafile=certifi.where())
def _is_cert_failure(err: BaseException) -> bool:
"""Is this a certificate-verification failure, however it is wrapped?
`urllib` does NOT let `ssl.SSLCertVerificationError` out. It catches it and
re-raises `urllib.error.URLError(err)`, carrying the original on `.reason`
- so `except ssl.SSLCertVerificationError` around `urlopen` matches
nothing, ever.
That is not a subtlety this file gets to record academically: the first
version of the fallback below was written exactly that way, shipped, and
failed on the machine it was written for with the very traceback it was
meant to prevent. The unit test passed throughout, because the fake it
used raised the bare SSL error - a shape real urllib never produces. A
stub that agrees with the author is worse than no test, and the test now
raises what urllib raises.
"""
reason = getattr(err, "reason", None)
return isinstance(err, ssl.SSLCertVerificationError) or \
isinstance(reason, ssl.SSLCertVerificationError)
def _urlopen(url: str, timeout: float = 60.0):
"""Open a URL, falling back to certifi's CA bundle on a verify failure.
Default first, certifi second, so the fix is additive: a machine whose
own store works keeps using it, and one with no store at all gets a
bundle rather than a traceback. certifi is already here - `requests` is a
hard dependency and brings it.
"""
try:
return urllib.request.urlopen(url, timeout=timeout)
except (urllib.error.URLError, ssl.SSLCertVerificationError) as err:
# Only a certificate problem is worth a second attempt. "No route to
# host" and "connection refused" arrive as URLError too, and retrying
# those with a different CA list changes nothing except how long the
# operator waits for the real message.
if not _is_cert_failure(err):
raise
ctx = _https_context()
if ctx is None:
raise
log.info("the system certificate store could not verify %s; "
"using the bundled CA list", url.split("/")[2])
return urllib.request.urlopen(url, timeout=timeout, context=ctx)
def _fetch(url: str, dest: Path, label: str) -> None:
"""Download with progress on stdout the supervisor can read.
@@ -56,7 +133,22 @@ def _fetch(url: str, dest: Path, label: str) -> None:
last = pct
log.info("download: %s %d%%", label, pct)
urllib.request.urlretrieve(url, dest, hook)
# Streamed rather than urlretrieve, only because urlretrieve offers no way
# to pass an SSL context and the whole point here is choosing one. The
# `download: <label> <n>%` lines are a contract: the supervisor parses
# them (`progressRe`) to put first-run progress in the tray, and without
# them a shop PC shows a stopped engine for five minutes after install.
with _urlopen(url) as resp:
total = int(resp.headers.get("Content-Length") or 0)
blocks, block_size = 0, 64 * 1024
with open(dest, "wb") as out:
while True:
chunk = resp.read(block_size)
if not chunk:
break
out.write(chunk)
blocks += 1
hook(blocks, block_size, total)
log.info("download: %s 100%%", label)
@@ -91,7 +183,7 @@ def setup_models(models_dir: Path) -> "list[str]":
import io
import zipfile
with urllib.request.urlopen(BUFFALO_SC_URL) as resp:
with _urlopen(BUFFALO_SC_URL) as resp:
payload = io.BytesIO(resp.read())
with zipfile.ZipFile(payload) as zf, \
zf.open("w600k_mbf.onnx") as src, \

View File

@@ -44,6 +44,9 @@ type App struct {
broker *agentmqtt.Client
stopBridge func()
hookURL string
// The resolved engine command, so engineMissing() and the supervisor are
// never looking at two different paths.
engineExe string
// Relays camera feeds to the webview so the engine's credential never has
// to travel in an <img> src, which a Chromium webview would strip anyway.
proxy *streamProxy
@@ -109,6 +112,9 @@ func (a *App) startup(ctx context.Context) {
if exe != "" && !filepath.IsAbs(exe) {
exe = filepath.Join(agentpaths.InstallRoot(), exe)
}
a.mu.Lock()
a.engineExe = exe
a.mu.Unlock()
logFile, _ := agentengine.LogFile(agentpaths.EngineLog())
a.sup = agentengine.New(agentengine.Options{
Command: func(c context.Context) *exec.Cmd {
@@ -154,13 +160,55 @@ func (a *App) startup(ctx context.Context) {
// not run yet, starting the supervisor would loop on a missing executable
// with nothing useful to say. The Start button still exists for the one
// case where somebody has deliberately stopped it.
if _, err := os.Stat(exe); err == nil {
if why := a.engineMissing(); why == "" {
a.sup.Start()
} else {
log.Printf("engine not installed yet (%s); run behavision-setup, then Start", exe)
log.Printf("%s (looked for %s)", why, exe)
}
}
// engineMissing says, in a sentence somebody can act on, why recognition
// cannot start here - or "" when it can.
//
// It exists because the answer was only ever given at startup, to a log file
// nobody on a shop counter opens. Pressing Start went straight to the
// supervisor, which reported what exec reported:
//
// engine failed to start: fork/exec /private/var/folders/c2/.../
// AppTranslocation/500A5354-.../d/Behavision.app/Contents/MacOS/engine/
// behavision: no such file or directory
//
// Every word of that is true and none of it says "run the setup tool". One
// function, consulted by the startup path, the Start button and the status
// panel, so the three cannot give three different accounts of one fact.
func (a *App) engineMissing() string {
a.mu.RLock()
exe := a.engineExe
a.mu.RUnlock()
if exe == "" {
return "The recognition engine is not set up on this computer yet."
}
if _, err := os.Stat(exe); err == nil {
return ""
}
msg := "The recognition engine is not installed on this computer yet. " +
"Run behavision-setup from the folder you unzipped, then press Start."
// macOS quarantines a downloaded app it cannot verify and runs it from a
// randomly named READ-ONLY copy - App Translocation. Every relative path
// then resolves inside that copy, which is why the engine folder appears
// to be missing from a bundle that plainly contains one, and why an
// install into it would not survive a restart. Detectable, unguessable,
// and fixed by one drag; saying nothing leaves somebody re-running a
// setup tool that cannot win.
if strings.Contains(exe, "/AppTranslocation/") {
msg = "macOS is running Behavision from a temporary read-only copy, " +
"because it was opened straight from Downloads. Move Behavision " +
"to your Applications folder and open it from there, then run " +
"behavision-setup."
}
return msg
}
// webhookURL is the loopback address the bridge is listening on, or empty
// before it has started.
func (a *App) webhookURL() string {
@@ -242,7 +290,7 @@ func (a *App) startPipeline(ctx context.Context) {
}
client, err := agentmqtt.NewClient(agentmqtt.ClientOptions{
BrokerURL: a.cfg.BrokerURL,
ClientID: "behavision-" + a.cfg.ClientID + "-" + a.cfg.SiteID,
ClientID: a.cfg.MQTTClientID(),
Username: a.cfg.BrokerUsername, Password: a.cfg.BrokerPassword,
CAFile: a.cfg.BrokerCAFile, Log: logger,
})
@@ -595,6 +643,11 @@ func (a *App) EngineStatus() EngineStatus {
if err != nil {
out.Error = err.Error()
}
// The supervisor's own error is an exec failure; this replaces it with
// the reason, which is the part that tells somebody what to do.
if why := a.engineMissing(); why != "" {
out.Error = why
}
ctx, cancel := context.WithTimeout(a.ctx, 4*time.Second)
defer cancel()
// A running process is not a working engine: on a memory-starved box the
@@ -609,6 +662,13 @@ func (a *App) EngineStatus() EngineStatus {
}
func (a *App) StartEngine() EngineStatus {
// Refused rather than attempted. Handing a missing path to the supervisor
// produces a retry loop and an exec error for a message.
if why := a.engineMissing(); why != "" {
st := a.EngineStatus()
st.Error = why
return st
}
if a.sup != nil {
a.sup.Start()
}

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View File

@@ -4,7 +4,7 @@
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>Behavision</title>
<script type="module" crossorigin src="./assets/index-6LYbNlbD.js"></script>
<script type="module" crossorigin src="./assets/index-DSNu_2FW.js"></script>
<link rel="stylesheet" crossorigin href="./assets/index-DOJ2bRrM.css">
</head>
<body>

View File

@@ -14,6 +14,18 @@ import { MAKES, makeById } from '../../../../shared/cameraMakes.js'
// signed off through.
const BLANK = { id: '', host: '', port: 554, path: '', username: '', password: '', max_width: 1280 }
// "This computer's own camera" - a webcam or a built-in FaceTime camera.
//
// The engine has supported it since the first version (`webcam: 0` picks a
// capture index instead of building an RTSP URL) and no screen has ever
// offered it: another case of the API being able to do something the UI
// could not reach. It matters most for the thing it was missing from, which
// is showing the product to somebody. A laptop's own camera gives real
// recognition, of real faces, in the room, depending on no network at all -
// where pointing a demo machine at a camera in another building depends on
// two internet connections and a tunnel staying up while you talk.
const WEBCAM = 'webcam'
export default function Cameras() {
const { data, error, reload } = usePolled(() => api.cameras(), 8000)
const [editing, setEditing] = useState(null)
@@ -190,7 +202,9 @@ function useStreamURLs(cams) {
function CameraSheet({ cam, onClose, onSaved }) {
const isNew = !cam.id
const [f, setF] = useState({ ...BLANK, ...cam, password: '', path: cam.path || (isNew ? MAKES[0].path : '') })
const [make, setMake] = useState(isNew ? MAKES[0].id : 'manual')
const [make, setMake] = useState(isNew ? MAKES[0].id : (cam.webcam != null ? WEBCAM : 'manual'))
const [index, setIndex] = useState(cam.webcam != null ? String(cam.webcam) : '0')
const local = make === WEBCAM
const [test, setTest] = useState(null)
const [busy, setBusy] = useState(null)
const [error, setError] = useState(null)
@@ -199,6 +213,7 @@ function CameraSheet({ cam, onClose, onSaved }) {
// Only overwrite the path when the preset has one, so choosing "I know the
// path" does not wipe what the installer already typed.
function chooseMake(e) {
if (e.target.value === WEBCAM) { setMake(WEBCAM); setTest(null); return }
const m = makeById(e.target.value)
setMake(m.id)
setF(prev => ({ ...prev, path: m.path || prev.path }))
@@ -212,6 +227,14 @@ function CameraSheet({ cam, onClose, onSaved }) {
if (v === '' || v === null || v === undefined) continue
out[k] = (k === 'port' || k === 'max_width') ? Number(v) : v
}
if (local) {
// An address and a webcam index are alternatives, not extras: the
// engine's source() takes the webcam first, so leaving a half-typed
// host behind would make the saved camera describe two different
// things and only one of them would be used.
for (const k of ['host', 'path', 'username', 'password']) delete out[k]
out.webcam = Number(index) || 0
}
return out
}
@@ -259,7 +282,7 @@ function CameraSheet({ cam, onClose, onSaved }) {
<p className="lead">Three things from the camera: its address, its make, and its password. Test before you save — a wrong address is the most common mistake.</p>
{error && <div className="err"><Icon.Warning size={15} />{error}</div>}
{isNew && (
{isNew && !local && (
<section className="formsection">
<h4>Find it</h4>
{scan === null && (
@@ -301,15 +324,20 @@ function CameraSheet({ cam, onClose, onSaved }) {
<em className="hint">Short, no spaces. It names this camera everywhere and cannot be changed later.</em>
</label>
)}
<div className="fieldrow">
<label className="field"><span>Address</span>
<input value={f.host} onChange={set('host')} placeholder="192.168.1.20" inputMode="decimal" />
<em className="hint">On a label on the camera, or in its own app under “network”.</em>
</label>
<label className="field narrow"><span>Port</span>
<input value={f.port} onChange={set('port')} inputMode="numeric" />
</label>
</div>
{local
? <label className="field narrow"><span>Camera number</span>
<input value={index} onChange={e => { setIndex(e.target.value); setTest(null) }} inputMode="numeric" />
<em className="hint">0 is the built-in camera. Try 1 if a second one is plugged in.</em>
</label>
: <div className="fieldrow">
<label className="field"><span>Address</span>
<input value={f.host} onChange={set('host')} placeholder="192.168.1.20" inputMode="decimal" />
<em className="hint">On a label on the camera, or in its own app under “network”.</em>
</label>
<label className="field narrow"><span>Port</span>
<input value={f.port} onChange={set('port')} inputMode="numeric" />
</label>
</div>}
</section>
<section className="formsection">
@@ -317,27 +345,34 @@ function CameraSheet({ cam, onClose, onSaved }) {
<label className="field"><span>Make of camera</span>
<select value={make} onChange={chooseMake}>
{MAKES.map(m => <option key={m.id} value={m.id}>{m.label}</option>)}
<option value={WEBCAM}>This computer’s own camera</option>
</select>
{chosen.note && <em className="hint">{chosen.note}</em>}
</label>
<label className="field"><span>Stream path</span>
<input className="mono" value={f.path} onChange={set('path')} placeholder="/Streaming/Channels/101" />
<em className="hint">Filled in from the make. Change it only if the camera’s own app says something else.</em>
{local
? <em className="hint">Recognition runs on this computer’s built-in or plugged-in camera. Nothing on the network is involved.</em>
: chosen.note && <em className="hint">{chosen.note}</em>}
</label>
{!local && (
<label className="field"><span>Stream path</span>
<input className="mono" value={f.path} onChange={set('path')} placeholder="/Streaming/Channels/101" />
<em className="hint">Filled in from the make. Change it only if the camera’s own app says something else.</em>
</label>
)}
</section>
<section className="formsection">
<h4>Sign-in to the camera</h4>
<div className="fieldrow">
<label className="field"><span>Username</span>
<input name="rtsp-account" autoComplete="off" value={f.username} onChange={set('username')} placeholder="admin" />
</label>
<label className="field"><span>Password</span>
<input type="password" name="rtsp-secret" autoComplete="new-password" value={f.password}
onChange={set('password')} placeholder={cam.has_password ? '(unchanged)' : ''} />
</label>
</div>
</section>
{!local && (
<section className="formsection">
<h4>Sign-in to the camera</h4>
<div className="fieldrow">
<label className="field"><span>Username</span>
<input name="rtsp-account" autoComplete="off" value={f.username} onChange={set('username')} placeholder="admin" />
</label>
<label className="field"><span>Password</span>
<input type="password" name="rtsp-secret" autoComplete="new-password" value={f.password}
onChange={set('password')} placeholder={cam.has_password ? '(unchanged)' : ''} />
</label>
</div>
</section>
)}
{test && (
test.ok

View File

@@ -1,11 +1,12 @@
[project]
name = "behavision"
version = "1.1.0"
dynamic = ["version"]
description = "Production face recognition over RTSP"
requires-python = ">=3.10"
dependencies = [
"numpy>=1.26,<2.0",
"opencv-python>=4.8.1",
# See requirements.txt for why numpy is uncapped and opencv is not.
"numpy>=1.26,<3.0",
"opencv-python>=4.8.1,<5",
"onnxruntime>=1.16",
"fastapi>=0.110",
"uvicorn>=0.29",
@@ -21,7 +22,11 @@ dependencies = [
]
[project.optional-dependencies]
dev = ["pytest>=8.0"]
# httpx is test-only and never ships in the wheel. The HTTP tests begin with
# `pytest.importorskip("httpx")` so a bare checkout still runs - which is
# right, and has a cost worth knowing: without it the suite reports 239 passed
# and quietly SKIPS 32 API tests. `pip install -e .[dev]` is how to get them.
dev = ["pytest>=8.0", "httpx>=0.27"]
[tool.setuptools.packages.find]
include = ["behavision*"]
@@ -31,3 +36,9 @@ behavision = ["static/*"]
[tool.pytest.ini_options]
testpaths = ["tests"]
# One version for the package and the wheel. Two literals in two files is how
# they came to disagree (1.0.0 here, 1.1.0 there) and how neither tracked a
# release.
[tool.setuptools.dynamic]
version = {attr = "behavision.__version__"}

View File

@@ -50,7 +50,28 @@ step "2. Agent and setup tool"
step "3. Engine source and wheel"
# The wheel is built with the checkout's own interpreter; requires-python is a
# statement about the SHOP PC, which setup enforces when it finds Python there.
# Stamp the tag into the wheel version. Without this every release built
# behavision-1.1.0-py3-none-any.whl, and `pip install --upgrade` on a machine
# that already had 1.1.0 is a no-op - so an engine fix reached nobody who had
# ever run setup, while the Go binaries beside it updated normally. Nothing
# reported a version either, so there was no way to tell a shop PC three
# releases behind from a current one.
#
# v0.5.6-demo -> 0.5.6+demo, which is valid PEP 440: a local segment takes
# alphanumerics and dots, never hyphens.
TMPVER=$(mktemp)
PEP440=$(printf '%s' "${TAG#v}" | sed 's/-/+/; s/[^0-9A-Za-z.+]/./g')
trap 'git checkout -- behavision/__init__.py 2>/dev/null || true; rm -f "$TMPVER"' EXIT
# A temp file rather than `sed -i`, whose argument handling differs between
# BSD and GNU - this script is run from a Mac today and that is not a reason
# to plant a portability trap in a release path.
sed "s/^__version__ = .*/__version__ = \"$PEP440\"/" behavision/__init__.py > "$TMPVER"
cat "$TMPVER" > behavision/__init__.py
.venv/bin/python -m pip wheel --no-deps --ignore-requires-python -q -w "$STAGE/engine-src" . 2>&1 | grep -v "DEPRECATION\|WARNING: Ignoring" || true
git checkout -- behavision/__init__.py
rm -f "$TMPVER"
trap - EXIT
echo " engine wheel version $PEP440"
ls "$STAGE"/engine-src/behavision-*.whl >/dev/null || { echo "wheel was not built" >&2; exit 1; }
cp pyproject.toml requirements.txt "$STAGE/engine-src/"
mkdir -p "$STAGE/engine-src/config" && cp config/default.yaml "$STAGE/engine-src/config/"

View File

@@ -1,5 +1,31 @@
numpy>=1.26,<2.0
opencv-python>=4.8.1
# numpy 2 is allowed, and that is what lets this install on a current Python.
# `<2.0` capped the resolver at numpy 1.26.4, whose newest wheel is cp312, so
# on a Mac with Python 3.14 pip fell back to BUILDING numpy from source and
# died in clang - two screens of C compiler output on a shop counter, for a
# version choice made silently by behavision-setup.
#
# Measured before changing it, nine runs of the detector guard per combination
# on one machine:
#
# numpy 1.26 / cv2 4.11 9 passed, 0 crashed
# numpy 2.0 / cv2 4.11 8 passed, 1 crashed
# numpy 1.26 / cv2 4.14 3 passed, 6 crashed
# numpy 2.0 / cv2 4.14 2 passed, 7 crashed
#
# numpy is not the variable there; OpenCV is. The crash was a test racing a
# shared cv2.FaceDetectorYN on purpose - undefined behaviour in C++, which 4.11
# usually turned into an exception and 4.14 usually turns into a segfault. The
# product never shares one (Engine._build_worker builds a detector per camera),
# and the test now runs that race out of process. With that fixed, the whole
# suite is 226 passed / 2 skipped on numpy 2.0.2, five runs out of five.
#
# opencv IS capped, and the two are not the same call. Everything above was
# measured on 4.x; OpenCV 5.0 is a major release this project has never run a
# real camera or an emotion model through, and an uncapped `>=4.8.1` means
# every NEW install silently gets it while every existing one keeps 4.11. Lift
# it after running a camera on 5.x, not before.
numpy>=1.26,<3.0
opencv-python>=4.8.1,<5
onnxruntime>=1.16
fastapi>=0.110
uvicorn>=0.29

View File

@@ -225,3 +225,42 @@ def test_an_inverted_per_camera_pair_is_rejected_not_stored(client):
tuning={"enroll_threshold": 0.8, "match_threshold": 0.5})
assert r.status_code == 400
assert c.get("/api/cameras").json() == []
# "This computer's own camera" — the demo case, and a capability the engine
# has always had with nothing able to reach it.
#
# It matters most for showing the product to somebody. A laptop's own camera
# gives real recognition, of real faces, in the room, depending on no network
# at all — where pointing a demo machine at a camera in another building
# depends on two internet connections and a tunnel staying up while you talk.
def test_this_computers_own_camera_can_be_added(client):
c, eng = client
r = c.post("/api/cameras", json={"id": "laptop", "webcam": 0})
assert r.status_code == 201, r.text
assert "laptop" in eng.workers
cam = next(x for x in c.get("/api/cameras").json() if x["id"] == "laptop")
# Returned, or the form cannot tell a webcam camera from a half-filled
# RTSP one when somebody opens it to edit.
assert cam["webcam"] == 0
assert eng.camera_store.get("laptop").source() == 0
def test_a_second_camera_index_is_kept(client):
"""0 is the built-in one; a plugged-in camera is usually 1. An index
silently coerced to 0 would open the wrong camera and look like the
setting had no effect."""
c, eng = client
assert c.post("/api/cameras", json={"id": "usb", "webcam": 1}).status_code == 201
assert eng.camera_store.get("usb").source() == 1
def test_a_webcam_camera_survives_a_reload(client, tmp_path):
"""It has to be on disk, not only in the running engine: a demo that
forgets its camera when the app restarts is worse than no demo."""
c, _ = client
c.post("/api/cameras", json={"id": "laptop", "webcam": 0})
again = CameraStore(tmp_path / "cameras.json").get("laptop")
assert again.source() == 0
assert again.safe_url() == "webcam:0"

View File

@@ -1,7 +1,32 @@
"""cv2.FaceDetectorYN caches its input size and is not thread-safe, so camera
workers must not share one. Skipped when the model is absent, matching the
faiss-optional pattern in test_index.py — the suite stays runnable with no
models installed."""
models installed.
The premise is proved in a SUBPROCESS, and that is the whole lesson of this
file. The first version raced a shared detector in-process and asserted that
an exception came back, because on OpenCV 4.11 one usually did. It is a data
race in C++: what it produces is undefined, and on 4.14 what it mostly
produces is a segmentation fault. Measured, nine runs each, same machine:
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
numpy is not the variable; OpenCV is, and 4.11 crashing once says the hazard
was always there and 4.11 merely survived it. A test that takes the whole
suite down two runs in three is worse than no test: it turns "we upgraded
OpenCV" into a CI failure with no failing assertion in it, which is the
hardest kind to read.
None of this reaches the product. `Engine._build_worker` constructs a
FaceDetector per camera, which is what the rule says and what the second test
here guards.
"""
import subprocess
import sys
import textwrap
import threading
from pathlib import Path
@@ -11,10 +36,16 @@ import pytest
from behavision.config import Config
from behavision.detection import YUNET_FILENAME, FaceDetector
MODELS = Path(__file__).resolve().parent.parent / "models"
ROOT = Path(__file__).resolve().parent.parent
MODELS = ROOT / "models"
pytestmark = pytest.mark.skipif(not (MODELS / YUNET_FILENAME).exists(),
reason="YuNet model not installed")
# A shared detector fails probabilistically, so one clean attempt proves
# nothing. Several do: the premise holds if ANY attempt misbehaves, and only
# an unbroken run of clean ones is evidence it has stopped being true.
ATTEMPTS = 6
def _detector():
d = Config().detection
@@ -44,11 +75,37 @@ def _race(det_a, det_b):
return errors
_SHARED_RACE = textwrap.dedent("""
import sys
sys.path.insert(0, {root!r})
from tests.test_detector_concurrency import _detector, _race
shared = _detector()
print("raced" if _race(shared, shared) else "clean")
""")
def _shared_race_outcome():
"""Run one shared-detector race out of process.
Returns "raced" (an exception came back), "crashed" (the process died,
which is the same premise arriving by a blunter route) or "clean".
"""
proc = subprocess.run([sys.executable, "-c", _SHARED_RACE.format(root=str(ROOT))],
capture_output=True, text=True, timeout=120)
if proc.returncode != 0:
return "crashed"
return proc.stdout.strip().splitlines()[-1] if proc.stdout.strip() else "clean"
def test_a_shared_detector_really_does_race():
"""Guards the premise: if this ever stops failing, the test below is
proving nothing and the per-camera split can be revisited."""
shared = _detector()
assert _race(shared, shared), "expected a shared detector to race"
seen = [_shared_race_outcome() for _ in range(ATTEMPTS)]
assert any(o != "clean" for o in seen), (
f"a shared detector survived {ATTEMPTS} races ({seen}) - if that is "
"reproducible, cv2.FaceDetectorYN may have become thread-safe and the "
"per-camera rule in CLAUDE.md can be revisited"
)
def test_per_camera_detectors_do_not_race():

View File

@@ -0,0 +1,164 @@
"""Downloading the models is the last step of every fresh install, and it ran
into the one macOS trap nothing else here does.
A python.org macOS build ships its own OpenSSL with NO trust store, and
populates one only when somebody double-clicks `Install Certificates.command`
in the Python folder. Measured on a colleague's Mac: the engine installed
perfectly - numpy, onnxruntime, faiss, all of it - and then could not fetch a
230 KB model file, ending setup in forty lines of traceback about `_ssl.c`.
"""
import io
import logging
import os
import ssl
import urllib.request
import pytest
from behavision import model_assets
class _Resp(io.BytesIO):
"""Enough of an http response for _fetch: read() and .headers."""
def __init__(self, payload: bytes):
super().__init__(payload)
self.headers = {"Content-Length": str(len(payload))}
def __enter__(self):
return self
def __exit__(self, *exc):
self.close()
return False
def _verify_error():
"""What urlopen ACTUALLY raises, which is not what it looks like.
urllib catches ssl.SSLCertVerificationError and re-raises
urllib.error.URLError(err), carrying the original on `.reason`. The first
version of this test raised the bare SSL error - a shape real urllib never
produces - so it passed against a fallback that could never fire, and the
fix shipped and failed on the machine it was written for with the exact
traceback it was meant to prevent.
"""
return urllib.error.URLError(ssl.SSLCertVerificationError(
"[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: "
"unable to get local issuer certificate"))
def test_a_machine_with_no_trust_store_falls_back_to_the_bundled_one(monkeypatch):
calls = []
def fake(url, timeout=None, context=None):
calls.append(context)
if context is None:
raise _verify_error()
return _Resp(b"ok")
monkeypatch.setattr(urllib.request, "urlopen", fake)
with model_assets._urlopen("https://example.invalid/m.onnx") as resp:
assert resp.read() == b"ok"
assert len(calls) == 2, f"expected a retry, got {calls}"
# Default FIRST, and that order is the point. On Windows and on a system
# Python the default context reads the machine's own certificate store,
# which is what makes a corporate proxy with its own root CA work.
# Replacing it unconditionally would break every site that has one in
# order to fix a different platform.
assert calls[0] is None
assert isinstance(calls[1], ssl.SSLContext)
def test_a_working_trust_store_is_used_as_is(monkeypatch):
calls = []
def fake(url, timeout=None, context=None):
calls.append(context)
return _Resp(b"ok")
monkeypatch.setattr(urllib.request, "urlopen", fake)
model_assets._urlopen("https://example.invalid/m.onnx").close()
assert calls == [None], "the machine's own certificate store was bypassed"
def test_a_real_network_failure_is_not_disguised_as_a_certificate_problem(monkeypatch):
"""A URLError is not automatically a certificate problem.
"No route to host" and "connection refused" arrive as URLError too, and
retrying those with a different CA list changes nothing except how long
the operator waits for the real message.
"""
calls = []
def fake(url, timeout=None, context=None):
calls.append(context)
raise urllib.error.URLError("no route to host")
monkeypatch.setattr(urllib.request, "urlopen", fake)
with pytest.raises(urllib.error.URLError):
model_assets._urlopen("https://example.invalid/m.onnx")
assert calls == [None], f"a dead network was retried as a CA problem: {calls}"
def test_progress_lines_survive_the_rewrite(tmp_path, monkeypatch, caplog):
"""`download: <label> <n>%` is a CONTRACT, not logging.
The supervisor parses it (progressRe) to put first-run progress in the
tray and the window, because the API is not up yet and a shop PC showing
a stopped engine for five minutes after install looks broken. Switching
off urlretrieve - needed because it offers no way to pass an SSL context -
is exactly the kind of change that drops it silently.
"""
payload = b"x" * (64 * 1024 * 8)
monkeypatch.setattr(urllib.request, "urlopen",
lambda url, timeout=None, context=None: _Resp(payload))
dest = tmp_path / "model.onnx"
with caplog.at_level(logging.INFO, logger="behavision.model_assets"):
model_assets._fetch("https://example.invalid/m.onnx", dest, "face detector")
assert dest.read_bytes() == payload
lines = [r.getMessage() for r in caplog.records]
pct = [ln for ln in lines if ln.startswith("download: face detector ")]
assert pct, f"no progress lines at all: {lines}"
assert "download: face detector 100%" in pct, pct
@pytest.mark.skipif(not os.environ.get("BEHAVISION_NETWORK_TESTS"),
reason="set BEHAVISION_NETWORK_TESTS=1 to reach github.com")
def test_against_a_real_machine_with_no_trust_store(tmp_path, monkeypatch):
"""The whole thing, over the real network, with the real failure.
Every stub above is a statement about what I believe urllib does, and the
first version of this file proved how much that is worth. Python's default
context honours SSL_CERT_FILE, so pointing it at an EMPTY file reproduces
a python.org macOS build exactly: a default context that trusts nobody.
certifi is loaded by path and is unaffected by the variable, so if the
fallback works here it works there.
"""
empty = tmp_path / "no-cas.pem"
empty.write_text("")
monkeypatch.setenv("SSL_CERT_FILE", str(empty))
monkeypatch.setenv("SSL_CERT_DIR", str(tmp_path / "nothing"))
# Not every interpreter can be put into that state, and pretending
# otherwise would turn this into a test that passes by not running. A
# Python linked against LibreSSL - which is what macOS Command Line Tools
# ships - reads the system keychain and ignores SSL_CERT_FILE entirely:
# measured, 128 CAs loaded with the variable pointing at an empty file. A
# python.org build links OpenSSL and honours it: 0 CAs, which is the
# condition being reproduced.
if ssl.create_default_context().get_ca_certs():
pytest.skip("this interpreter ignores SSL_CERT_FILE (%s), so the "
"no-trust-store condition cannot be reproduced here"
% ssl.OPENSSL_VERSION)
# The premise: the default context really is broken in this process.
with pytest.raises(urllib.error.URLError) as caught:
urllib.request.urlopen(model_assets.YUNET_URL, timeout=30)
assert model_assets._is_cert_failure(caught.value), caught.value
dest = tmp_path / "yunet.onnx"
model_assets._fetch(model_assets.YUNET_URL, dest, "face detector")
assert dest.stat().st_size > 200_000

View File

@@ -0,0 +1,81 @@
"""Why a camera "will not connect", when the real answer is that the computer
asking is in the wrong building.
Asked directly by the owner, about his own cameras, from his phone's
connection: *"when i connect from my mobile internet the cameras wont connect,
why is that"*. The answer was `cannot reach 192.168.1.121:554 - Operation
timed out`, which reads as a broken camera and sends somebody to re-type an
address and a password that were always correct.
"""
import pytest
from behavision import capture
@pytest.fixture
def on(monkeypatch):
def _set(ip):
monkeypatch.setattr(capture, "_local_ipv4", lambda: ip)
return _set
def test_a_different_network_says_so_and_says_what_to_do(on):
on("10.11.12.13") # a phone's tethered network
hint = capture._wrong_network_hint("192.168.1.121")
assert "not the camera's network" in hint
# The action, not just the diagnosis: no setting on this screen fixes it.
assert "computer in the shop" in hint
def test_the_same_network_sends_you_to_the_camera_instead(on):
on("192.168.1.120")
hint = capture._wrong_network_hint("192.168.1.121")
assert "on that network" in hint
assert "powered on" in hint
# Two states that need opposite actions must not share a sentence.
assert "not the camera's network" not in hint
@pytest.mark.parametrize("host", [
"8.8.8.8", # plainly routable
"203.0.113.9", # TEST-NET-3: `is_private` calls this private, and it is
# not a LAN address - which is why the check spells out
# the RFC1918 blocks instead of asking is_private
"100.64.0.5", # carrier-grade NAT, what a mobile network hands out
])
def test_an_address_that_is_not_a_lan_address_gets_no_hint(on, host):
"""A routable address unreachable from here is an ordinary network fault,
and inventing a story about private networks would be wrong."""
on("192.168.1.120")
assert capture._wrong_network_hint(host) == ""
def test_a_name_gets_no_hint(on):
"""Nothing can be concluded about `camera.local` from the string, and a
guess here is a confident wrong answer in the place people look first."""
on("192.168.1.120")
assert capture._wrong_network_hint("camera.local") == ""
def test_loopback_gets_no_hint(on):
on("192.168.1.120")
assert capture._wrong_network_hint("127.0.0.1") == ""
def test_with_no_network_at_all_it_still_names_the_cause(on):
"""A machine with no route cannot say which network it is on, and must not
pretend: the private-address fact is still true and still the reason."""
on("")
hint = capture._wrong_network_hint("192.168.1.121")
assert "private address" in hint
assert "this computer is on" not in hint
def test_the_hint_reaches_the_message_a_person_reads():
"""The whole point is the sentence on the screen, not a helper nobody
calls. Port 1 on a private address refuses or times out immediately."""
ok, msg = capture._tcp_reachable("rtsp://192.168.1.121:1/ch0", 1.0)
assert not ok
assert "192.168.1.121" in msg
# Whichever branch the OS takes, the explanation travels with it.
assert "network" in msg