9 Commits

Author SHA1 Message Date
c6a2c392d9 Paste the camera's RTSP address instead of taking it apart by hand
Asked directly: "our cameras have an rtsp url, we can use that to connect them
to this software right". Yes - and that has always been the mechanism, which
is the point. CameraConfig.source() builds exactly that URL from the parts,
and CameraConfig.url has always accepted a whole one and taken priority over
them. No form ever offered it.

So an operator holding the address their camera's own app shows had to split
it into five fields by eye. That is where a password containing @ or / goes
wrong, and this repository has already been bitten once by unencoded @ in RTSP
credentials.

parseRtspUrl lives in shared/cameraMakes.js and is imported by BOTH forms, for
the same reason the make picker is: two copies would be worse than not
offering it, because an operator trusts a filled-in field. A test asserts both
import it.

Decisions worth keeping:

- Split into fields, not stored whole. Everything else on the form - Test, the
  make picker, editing later, and the rule that a password is never returned
  to the browser - works on the parts. A URL kept intact would carry the
  password back out to every screen that reads a camera.
- WHATWG splits user info at the LAST @, which is what makes an unencoded @
  inside a password parse the way a person means it. An operator doing it by
  eye would put "p" in the password box and "ssw0rd@192.168.1.121" in the
  address box.
- Percent-encoded credentials are DECODED, because source() encodes again
  when it rebuilds the URL. Keeping them encoded would double-encode and the
  camera would refuse a password that is correct.
- The scheme is optional, structure is not. Without requiring a slash,
  "nonsense" parses as a perfectly good hostname and silently fills the
  Address field with it - a wrong answer that looks like it worked. A bare
  address is refused too: the Address field already takes one.
- A query string stays with the path. Some cameras carry the channel there,
  and dropping it opens the wrong channel - which looks like a camera pointed
  somewhere unexpected.
- A URL carrying no credentials does not wipe a password already typed.

Tested through node from pytest, the same pattern test_dashboard.py uses, and
skipped when node is absent so the suite stays dependency-light.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 18:29:59 +05:30
b296e8a74a A demo does not need a tunnel; it needs the laptop's own camera
Asked after the mobile-internet question: could a VPN let the office cameras
be shown in the demo. Three jobs get confused there and only one needs one.

Seeing the estate from anywhere already works and needs nothing - viewer mode
plus LiveHub is exactly that, outbound, no installation and no credential.

Demonstrating recognition is better done on the demo machine's own camera.
`webcam: 0` picks a capture index instead of building an RTSP URL and the
engine has supported it since the first version: CameraStore round-trips it,
source() returns the index, safe_url() reports webcam:0, and
POST /api/cameras {"id":"laptop","webcam":0} has always worked. No screen
offered it - the same gap this repo already records for the customer record
and per-camera tuning. It is now an option in the make picker, and it is the
strongest demo available: real faces, in the room, depending on no network.
A demo pointed at a camera in another building depends on two internet
connections and a tunnel staying up while somebody is talking.

The address and the index are alternatives, not extras: source() takes the
webcam first, so a half-typed host left behind would make the saved camera
describe two things and use one. The scan, the path and the camera password
are hidden for a local camera because none of them mean anything.

A tunnel is still right for one case - running the engine on a remote machine
against the office's own cameras - and still wrong for the product: it is
per-site infrastructure on every shop PC, and it gives head office
network-level access into a customer's LAN, where today we can read a
camera's picture and nothing else.

Also: installing httpx took the suite from 239 passed to 271. The HTTP tests
importorskip it so a bare checkout runs, which means the number at the bottom
of a run is not the number of tests that exist. Added to the dev extra.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 18:18:57 +05:30
c9be9b5807 "The cameras won't connect from my mobile internet" was answered by a timeout
Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics and the product was not giving it.

A camera lives on the shop's LAN behind a router. 192.168.1.121 means
"something on the network I am attached to" and nothing more - from mobile
data, a hotel or head office it resolves to nobody, or to a completely
different device holding that number. There is no route in from the internet
and there must not be: an RTSP camera reachable from outside is how a shop's
cameras end up being watched by strangers.

That is why the product is split the way it is - the shop PC is the only
machine on the camera's LAN, and every other surface reaches it outbound,
which is what makes Watch live work from anywhere while nothing connects in.

What was wrong is the message. "cannot reach 192.168.1.121:554 - Operation
timed out" reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. _wrong_network_hint names the cause
and separates two states that need opposite actions:

  on that network   -> check the camera is powered on and the address is right
  somewhere else    -> the COMPUTER is in the wrong place; no setting fixes it

- The local address comes from a connected UDP socket that sends nothing. It
  only fixes a route so the kernel will name the source address.
- The LAN ranges are spelled out, not is_private. That property also covers
  carrier-grade NAT and the documentation networks, and telling somebody who
  typed 203.0.113.9 that it is "on the shop's own network" is a confident
  wrong answer in the place people look first. Found by a test using that
  address as its example of a PUBLIC one.
- A DNS name gets no hint: nothing can be concluded about camera.local from
  the string, and guessing is the failure mode this message exists to fix.
- With no network at all it still names the cause and drops the comparison,
  rather than claiming to know which network this machine is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 18:00:44 +05:30
0558344dc2 urllib does not let SSLCertVerificationError out, and a fake said it did
The certificate fix shipped and failed on the machine it was written for, with
the exact traceback it was meant to prevent. The retry was written

    except ssl.SSLCertVerificationError:

and urllib never raises that from urlopen. It catches it and re-raises
urllib.error.URLError(err), carrying the original on .reason. So the except
matched nothing, ever, and the fallback could not fire.

The unit test passed throughout, because the stub it used raised the bare SSL
error - a shape real urllib never produces. That is the lesson: a fake that
agrees with the author is worse than no test, because it converts an untested
path into a tested-looking one. This file already says that about
UPDATE ... RETURNING and about the in-memory API fake, and it got written
again anyway.

_is_cert_failure checks the exception and its .reason, and the tests now raise
URLError(SSLCertVerificationError(...)) - what the traceback actually shows. A
plain URLError is re-raised untouched, and a test asserts no second attempt is
made for one.

Beside the stubs there is now a real reproduction, opt-in behind
BEHAVISION_NETWORK_TESTS=1. Python's default context honours SSL_CERT_FILE, so
an empty file gives a context that trusts nobody - the python.org condition
exactly - while certifi is loaded by path and is unaffected. It skips rather
than passes where it cannot reproduce that, and the difference is measured:

    macOS Command Line Tools  LibreSSL 2.8.3   128 CAs with an empty CA file
    python.org / pyenv build  OpenSSL 3.5.8      0 CAs -> reproduces it

Checked for teeth by putting the shipped except back: both the corrected stub
test and the live one fail, and pass again when it is restored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:36:53 +05:30
8f07dee048 The engine stopped updating, and said "installed" every time
The SSL fix shipped and did not reach the machine it was written for. That
log said so, one line above the tick:

  behavision is already installed with the same version as the provided
  wheel. Use --force-reinstall to force an installation of the wheel.
   [ok] Engine and dependencies installed

and the traceback below it still pointed at model_assets.py line 59,
urllib.request.urlretrieve - code the fix had deleted.

Two frozen literals caused it: version = "1.1.0" in pyproject.toml and
__version__ = "1.0.0" in behavision/__init__.py. They disagreed with each
other and neither tracked a release, so every release built
behavision-1.1.0-py3-none-any.whl and pip install --upgrade on a machine that
already had 1.1.0 is a no-op. The comment beside that call claimed the
opposite.

The shape of the damage is what makes it bad. The Go binaries - app, agent,
setup tool - are rebuilt every release and updated normally. So a shop PC ran
a current app supervising an engine several releases old, and nothing said
which: /api/health reported the model, the paths, the cameras and the gallery,
and no version at all.

- One version, in the package, read by pyproject through
  [tool.setuptools.dynamic]. In a checkout it reads 0.0.0+dev: a
  plausible-looking number on a developer's /api/health is worse than none.
- pip install --force-reinstall --no-deps <wheel>, after the ordinary
  --upgrade. --upgrade settles the dependencies; the second call guarantees our
  own code is the code in the folder. --no-deps keeps it cheap - forcing the
  dependencies too would re-download ~300 MB every run. A rebuild at an
  unchanged version is the ordinary case while developing, so this must not
  rely on the version moving.
- release.sh stamps the tag: v0.5.6-demo -> 0.5.6+demo, valid PEP 440. Into a
  copy of the line, reverted in a trap, so the tree is never left dirty.

/api/health reports version now. Without it there is no way to tell a shop PC
three releases behind from a current one, which is how this survived several
releases.

Reproduced end to end before fixing, against real wheels on Python 3.12: two
builds of the same version, --upgrade leaves the old code in place and prints
the same sentence the colleague's Mac printed, --force-reinstall --no-deps
replaces it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:30:24 +05:30
3cddd9c2e1 The install finished and then could not download a 230 KB file
With the version ceiling and the widened numpy pin in place, setup succeeded
on the Mac that found them - Python 3.14 chosen and accepted, numpy 2.5.3,
onnxruntime 1.30, faiss 1.15.1, the engine itself - and died on the last step,
fetching the YuNet model:

  ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
  certificate verify failed: unable to get local issuer certificate

A python.org macOS build ships its own OpenSSL with NO trust store, and
populates one only when somebody double-clicks Install Certificates.command in
the Python folder. Nobody installing face-recognition software has a reason to
know that exists, and the failure is forty lines of traceback about _ssl.c at
the end of a ten-minute install.

_urlopen tries the default context first and retries with certifi's bundle on
a verification failure. The order is the design:

- Default first, because on Windows and on a system or Homebrew Python the
  default context reads the machine's own certificate store, which is what
  makes a corporate proxy with its own root CA work. Replacing it
  unconditionally would break every site that has one to fix a different
  platform.
- certifi second, because it is already installed: requests is a hard
  dependency and brings it.
- URLError is re-raised untouched. "No route to host" and "no trust store" are
  different problems, and retrying the first with a different CA list only
  delays the real message.

urlretrieve had to go, since it offers no way to pass a context - exactly the
kind of rewrite that silently drops something. The `download: <label> <n>%`
lines are a contract: supervisor.go's progressRe parses them to put first-run
progress in the tray, because the API is not up yet and a shop PC showing a
stopped engine for five minutes looks broken. A test asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:21:41 +05:30
248025cdf9 The Python ceiling was defeated by the venv the failure left behind
makeVenv reused any environment already on disk, whatever Python built it.
The machine that found the version bug already had a runtime built by 3.14,
left there by the run that failed - so with the ceiling in place setup would
choose a good interpreter, reach makeVenv, find the 3.14 environment, keep it,
and die in the same clang error as before.

A fix a user cannot reach because the bug's own debris is in the way is not a
fix, and it would have read as the release not working.

It now asks the interpreter inside an existing environment what it is and
rebuilds when the answer is unsupported, saying so. Rebuilding costs a
re-download of the libraries and nothing else - the models live in the state
root. An environment that cannot be asked counts as unusable too: a
half-created one answers nothing, and reusing it fails later in pip with an
error about a package rather than about the environment.

Tested against real environments rather than a fake, because what is under
test is what an interpreter on disk reports about itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:11:47 +05:30
ff4f95c3b0 Four silent failures a demo on somebody else's Mac walked straight into
Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.

## behavision-setup chose the Python least likely to work

findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".

maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.

## numpy<2.0 was the cap; OpenCV was the hazard

Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:

  numpy 1.26 / cv2 4.11    9 passed, 0 crashed
  numpy 2.0  / cv2 4.11    8 passed, 1 crashed
  numpy 1.26 / cv2 4.14    3 passed, 6 crashed
  numpy 2.0  / cv2 4.14    2 passed, 7 crashed

numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.

It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.

opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.

## One MQTT client id for a whole shop, so two PCs fought over it

behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:

  broker connected / broker connection lost: EOF / broker connected / EOF ...

The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.

MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.

## "no such file or directory" for an engine nobody had installed

Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.

engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:09:25 +05:30
48a30d97db Live camera view in the app, and the green light that was lying about it
Two changes, and the second was found by verifying the first.

## Watching a camera from the app, in another building

Snapshots answer "is that camera working". They do not answer "what is
happening in my shop right now", which is what somebody who opens the app away
from the counter is asking. Head office's browser already had that answer -
LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is
watching and pushes JPEG frames for as long as somebody is - and the app could
not reach it.

cloud.CameraLive opens that feed and the app's own loopback relay re-emits it
as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img>
cannot render that, and an <img> renders MJPEG natively - so a tile is an
ordinary <img> pointed at loopback whether the camera is in this room or
another city.

- Reconnecting happens in the relay, not the page. The server caps one push at
  five minutes, so doing it here means the <img> never sees the stream end.
- The headers are flushed before the first frame. Go writes them on the first
  body write, so without that the whole response waits for the shop PC to
  start pushing. Measured against production: 30 seconds and not even a
  Content-Type, which surfaces as the request timing out.
- One camera at a time. Watching makes a shop PC upload, so a grid that went
  live at once would put an estate's worth of cameras on the wire because
  somebody opened a page.
- live.mjpeg is behind the same per-run token as the engine routes, and a
  wrong token is a 404 that never reaches head office at all.
- CameraLive uses its own HTTP client: the shared one's 30s timeout covers the
  whole response and would sever a working view every thirty seconds - the
  trap that made the server set WriteTimeout to zero for its own SSE endpoint.

## A camera read "Connected" for 34 minutes after the shop PC went blind

Which is why the verification above looked like a failure: head office
registered the viewer and no frame ever came.

reportWith returns early when the engine is unreachable - correctly, it has
nothing to say - so the last state it sent stays in the database looking
current. Measured live: cam2 and entrance both reading Connected, in green,
with last_seen_at 34 minutes old, while the heartbeat from the same PC said
cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing.

false could not be the answer. It means "this camera is not connecting", which
sends an installer to check cabling on a camera that was working perfectly the
last time anybody could ask it. So there are four states and one function:

  connected       reported recently, and working
  not_connecting  reported recently, and the stream will not open
  waiting         no shop PC has ever reported this camera
  stale           reported once, and not lately

- Connected is CLEARED when stale or waiting. A stale true left in place stays
  available to every client reading the field directly, and leaves two fields
  on one object disagreeing - how the shops screen once came out labelled
  Working, in green, above "2 of 3 cameras not connecting".
- Computed in scanCamera, so every camera anybody reads passes through it. A
  state computed per handler is one a handler forgets, and this had already
  reached three screens.
- CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning
  as three missed heartbeats - an indicator that cries wolf gets ignored.
- An unparseable last_seen_at is stale. It should be impossible, which is why
  it must not fall through to the state that says everything is fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 16:48:43 +05:30
40 changed files with 2715 additions and 145 deletions

443
CLAUDE.md
View File

@@ -3359,3 +3359,446 @@ live one.
- **With no engine AND nobody signed in, the engine error is still the answer.**
There is nothing else to show and the person is most likely setting this PC
up; naming head office there points them at a step they have not reached.
## Watching a camera from the app, in another building
Snapshots answer *"is that camera working"*. They do not answer *"what is
happening in my shop right now"*, which is what somebody who opens the app
away from the counter is asking. Head office's browser already had the answer
— `LiveHub` plus `cameras.Live`, where the shop PC asks outbound whether
anybody is watching and pushes JPEG frames up for exactly as long as somebody
is — and the app could not reach it.
`cloud.CameraLive` opens that feed and the app's own loopback relay re-emits
it as **multipart MJPEG**, which is the whole trick: frames arrive base64 over
SSE, an `<img>` cannot render that, and an `<img>` renders MJPEG natively. So a
tile is an ordinary `<img>` pointed at loopback whether the camera is in this
room or another city, and no screen has to know which.
- **Reconnecting happens in the relay, not the page.** The server caps one push
at five minutes so a tab left open for a week cannot leave a shop uploading
for a week. Doing it here means the `<img>` never sees the stream end.
- **The headers are flushed before the first frame.** Go writes them on the
first body write, so without that the whole response — status line included —
waits for the shop PC to start pushing. Measured against production: thirty
seconds and not even a `Content-Type`, which surfaces as the *request* timing
out rather than a stream that has not painted yet.
- **One camera at a time.** Watching makes a shop PC upload, so a grid that
went live at once would put an estate's worth of cameras on the wire because
somebody opened a page. `Watch live` is per tile and toggles the previous one
off.
- **`live.mjpeg` is behind the same per-run token as the engine routes**, and a
wrong token is a 404 that never reaches head office at all. It is a live view
of a shop floor; the relay being on loopback is not on its own a control.
- **`CameraLive` uses its own HTTP client.** The shared one has a 30-second
timeout that covers the whole response and would therefore sever a working
live view every thirty seconds — the same trap that made the server set
`WriteTimeout` to zero for its own SSE endpoint.
## A camera read "Connected" for 34 minutes after the shop PC went blind
Found while verifying the live view against production, and it is the reason
that verification looked like a failure: head office registered the viewer and
no frame ever came.
`reportWith` returns early when the engine is unreachable — correctly, because
it has nothing to say — so the last state it sent **stays in the database
looking current**. Measured on the live estate: `cam2` and `entrance` both
reading **Connected**, in green, with `last_seen_at` thirty-four minutes old,
while the heartbeat from the same PC said `cameras_up: 0, cameras_total: 0`.
Two surfaces reading two stored fields and disagreeing about one fact.
`false` could not be the answer. It means *"this camera is not connecting"*,
which sends an installer to check cabling on a camera that was working
perfectly the last time anybody could ask it. So there are four states, not
three, and `api.CameraState` is the one function that decides them:
| state | meaning | what to do |
|---|---|---|
| `connected` | reported within `CameraStaleAfter`, and working | — |
| `not_connecting` | reported recently, and the stream will not open | check the address, password, cabling |
| `waiting` | no shop PC has ever reported this camera | it has not reached the PC yet |
| `stale` | reported once, and not lately | check the PC is on and Behavision is running |
- **`Connected` is CLEARED when the state is `stale` or `waiting`.** Leaving a
stale `true` in place keeps the lie available to every client that reads the
field directly — a mobile app, a script, an older desktop build — and leaves
two fields on one object disagreeing, which is exactly how the shops screen
once came out labelled **Working**, in green, above *"2 of 3 cameras not
connecting"*.
- **It is computed in `scanCamera`**, so every camera anybody reads passes
through it. A state computed per handler is a state one handler forgets, and
this one had already reached three screens.
- **`CameraStaleAfter` is 5 minutes — five missed reports, not one.** The agent
reports on a 60-second tick, so one miss is a dropped packet. Same reasoning
as a site being offline after three missed heartbeats: an indicator that
cries wolf is one people learn to ignore.
- **An unparseable `last_seen_at` is stale**, not connected. It should be
impossible, which is precisely why it must not fall through to the state that
says everything is fine.
## A demo on somebody else's Mac found four things, all of them silent
Three failures in one afternoon on a colleague's machine, plus one the fixing
uncovered. Every one produced a message that was true and useless.
### behavision-setup chose the Python least likely to work
`findPython` walked `3.14, 3.13, 3.12, 3.11, 3.10` and took the first hit — a
floor with **no ceiling**, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels for anything. It picked
3.14, pip found no numpy wheel for cp314 (`numpy<2.0` caps the resolver at
1.26.4, whose newest is cp312), fell back to building numpy from source and
produced `ERROR: Unknown compiler(s)`; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
`<arm_neon.h> is intended only for ARM and AArch64 targets`.
Two screens of C compiler output on a shop counter, for a version choice this
program made silently. `maxMinor` refuses in one line before anything is
downloaded, and **"too new" is a different message from "too old"** — telling
somebody holding Python 3.14 that no Python was found sends them to install a
newer one, which is the direction that just failed. It is a *wheel-availability*
ceiling, not a language one: onnxruntime is the binding dependency today
(cp314 is its newest), numpy publishes further ahead, and opencv ships a
stable-ABI wheel that covers everything.
### `numpy<2.0` was the cap; OpenCV was the hazard
Widening to `<3.0` needed proof, and the proof found something else. Nine runs
of the detector guard per combination, one machine, one sitting:
```
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
```
**numpy is not the variable; OpenCV is** — the third row is numpy 1.26. The
crash was `test_a_shared_detector_really_does_race`, which races a shared
`cv2.FaceDetectorYN` on purpose to prove the per-camera rule. That is undefined
behaviour in C++: 4.11 usually turned it into an exception, 4.14 usually turns
it into a **segfault**, and 4.11 crashing once says the hazard was always there
and 4.11 merely survived it.
It never reached the product — `Engine._build_worker` builds a detector per
camera, which is the rule and is what the second test guards. What it reached
was the suite: two runs in three died with **no failing assertion in them**,
turning "we upgraded OpenCV" into the hardest kind of CI failure to read. The
race now runs in a **subprocess**, so a segfault is an observed outcome rather
than the end of the run, and one clean attempt proves nothing — the premise
holds if *any* of several attempts misbehaves. With that fixed the suite is
226 passed / 2 skipped on numpy 2.0.2, five runs out of five.
`opencv-python` stays capped below 5. Everything above was measured on 4.x, and
an uncapped `>=4.8.1` means every NEW install silently gets a major release
this project has never run a real camera through while every existing one keeps
4.11.
### And the fix was defeated by the wreckage of the bug
`makeVenv` reused any environment already on disk, whatever Python built it.
That machine had a runtime built by **3.14**, left behind by the run that
failed — so with the ceiling in place setup would choose a good interpreter,
reach `makeVenv`, find the 3.14 environment, keep it, and die in the same clang
error as before. A fix a user cannot reach because the bug's own debris is in
the way is not a fix, and it would have read as the release not working.
It now asks the interpreter inside an existing environment what it is and
rebuilds when the answer is unsupported, saying so. Rebuilding costs a
re-download of the libraries and nothing else — the models live in the state
root, not in there. An environment that cannot be asked counts as unusable
too: a half-created one answers nothing, and reusing it fails later in pip
with an error about a package rather than about the environment.
### One MQTT client id for a whole shop, so two PCs fought over it
`behavision-<client>-<site>` is the same string on every computer claimed to
one site. MQTT requires client ids to be unique and a broker enforces it by
disconnecting the older session when a new one arrives with the same id, so the
colleague's Mac and the shop's own till took turns kicking each other off:
```
broker connected / broker connection lost: EOF / broker connected / EOF / ...
```
**The damage is not confined to the new machine.** The shop's till is the other
half of that loop, so somebody signing in on a laptop to look at the product
stops a live shop delivering visits — and from each end it reads as an unstable
network, because nothing says otherwise.
`Config.MQTTClientID()` appends a per-installation id, minted on first load and
written back so an existing install gets one without anybody doing anything.
The site stays in the name because that is what a broker log is read *by*. A
config that could not be written falls back to a per-run id rather than a
shared one: the right failure is a new name in the log after a restart, not the
collision this exists to end.
### "no such file or directory" for an engine nobody had installed
Pressing Start with no engine went straight to the supervisor, which reported
what `exec` reported:
```
engine failed to start: fork/exec /private/var/folders/c2/.../AppTranslocation/
500A5354-.../d/Behavision.app/Contents/MacOS/engine/behavision:
no such file or directory
```
Every word true, none of it saying *run the setup tool*. The startup path did
have that sentence — in a log file nobody on a shop counter opens.
`App.engineMissing()` is now the one function the startup path, the Start
button and the status panel all consult, so three surfaces cannot give three
accounts of one fact.
It also names **App Translocation**, which is in that path and is unguessable.
macOS quarantines a downloaded app it cannot verify and runs it from a randomly
named read-only copy, so every relative path resolves inside that copy — which
is why the engine folder appears missing from a bundle that plainly contains
one, and why installing into it would not survive a restart. Fixed by dragging
the app to Applications; saying nothing leaves somebody re-running a setup tool
that cannot win. The product is unsigned, so this is the *normal* first-run
state on every Mac, not an edge case.
### And then it could not download a 230 KB file
With all of the above fixed the install succeeded on that Mac - Python 3.14
chosen and accepted, numpy 2.5.3, onnxruntime 1.30, faiss 1.15.1, the engine
itself - and setup died on the last step, fetching the YuNet model:
```
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
certificate verify failed: unable to get local issuer certificate
```
A python.org macOS build ships its **own OpenSSL with no trust store**, and
populates one only when somebody double-clicks `Install Certificates.command`
in the Python folder. Nobody installing face-recognition software has any
reason to know that exists, and the failure is forty lines of traceback about
`_ssl.c` at the end of a ten-minute install.
`_urlopen` tries the default context first and retries with **certifi's**
bundle on a verification failure. The order is the whole design:
- Default first, because on Windows and on a system or Homebrew Python the
default context reads the machine's own certificate store - which is what
makes a corporate proxy with its own root CA work. Replacing it
unconditionally would break every site that has one in order to fix a
different platform.
- certifi second, because it is already installed: `requests` is a hard
dependency and brings it.
- A `URLError` is re-raised untouched. "No route to host" and "no trust store"
are different problems, and retrying the first with a different CA list only
delays the real message.
`urlretrieve` had to go, since it offers no way to pass a context - and that is
exactly the kind of rewrite that silently drops something. The
`download: <label> <n>%` lines are a **contract**: `supervisor.go`'s
`progressRe` parses them to put first-run progress in the tray, because the API
is not up yet and a shop PC showing a stopped engine for five minutes after
install looks broken. `tests/test_model_download.py` asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.
## The engine stopped updating, and said "installed" every time
The SSL fix above was released and **did not reach the machine it was written
for**. Its log said so, one line above the tick:
```
behavision is already installed with the same version as the provided wheel.
Use --force-reinstall to force an installation of the wheel.
[ok] Engine and dependencies installed
```
and the traceback that followed still pointed at `model_assets.py", line 59,
in _fetch / urllib.request.urlretrieve` — code the fix had deleted.
Two frozen literals caused it: `version = "1.1.0"` in `pyproject.toml` and
`__version__ = "1.0.0"` in `behavision/__init__.py`. They disagreed with each
other and neither tracked a release, so **every release built
`behavision-1.1.0-py3-none-any.whl`** and `pip install --upgrade` on a machine
that already had 1.1.0 is a no-op. The comment beside that call even claimed
the opposite: *"`--upgrade` so re-running after a new release replaces the
engine rather than leaving the old one in place and reporting success."*
The shape of the damage is what makes it bad. The Go binaries — the app, the
agent, the setup tool — are rebuilt every release and updated normally. So a
shop PC ran a current app supervising an engine several releases old, and
nothing anywhere said which: `/api/health` reported the model, the paths, the
cameras and the gallery, and **no version at all**.
Three changes, and the second is the one that does not depend on the first
being remembered:
- **One version, in the package**, read by `pyproject.toml` through
`[tool.setuptools.dynamic]`. Two literals in two files is how they came to
disagree. In a checkout it reads `0.0.0+dev`, because a plausible-looking
number on a developer's `/api/health` would be worse than none.
- **`pip install --force-reinstall --no-deps <wheel>`, after the ordinary
`--upgrade`.** `--upgrade` settles the dependencies; the second call
guarantees our own code is the code in the folder. `--no-deps` is what keeps
it cheap — forcing the dependencies too would re-download ~300 MB of numpy,
OpenCV and onnxruntime on every run. A rebuild at an unchanged version is the
ordinary case while developing, so this must not rely on the version moving.
- **`release.sh` stamps the tag**: `v0.5.6-demo` → `0.5.6+demo`, valid PEP 440
(a local segment takes alphanumerics and dots, never hyphens). Into a copy of
the line, reverted in a `trap`, so the tree is never left dirty.
`/api/health` reports `version` now. Without it there is no way to tell a shop
PC three releases behind from a current one, which is precisely how this
survived several releases.
Reproduced end to end before fixing, against real wheels on Python 3.12: two
builds of the same version, `--upgrade` leaves the old code in place and prints
the same sentence the colleague's Mac printed, `--force-reinstall --no-deps`
replaces it.
### `urllib` does not let `SSLCertVerificationError` out, and a fake said it did
The certificate fix above shipped and **failed on the machine it was written
for, with the exact traceback it was meant to prevent**. The retry was written
```python
except ssl.SSLCertVerificationError:
```
and urllib never raises that from `urlopen`. It catches it and re-raises
`urllib.error.URLError(err)`, carrying the original on `.reason`. So the
`except` matched nothing, ever, and the fallback could not fire.
The unit test passed throughout, because the stub it used raised the bare SSL
error — **a shape real urllib never produces**. That is the whole lesson: a
fake that agrees with the author is worse than no test, because it converts an
untested path into a tested-looking one. The same sentence is already in this
file about `UPDATE ... RETURNING` and about the in-memory API fake, and it was
written again here anyway.
`_is_cert_failure` checks the exception **and** its `.reason`, and the tests
now raise `URLError(SSLCertVerificationError(...))`, which is what his
traceback shows. A plain `URLError` — "no route to host" — is re-raised
untouched: retrying that with a different CA list changes nothing except how
long the operator waits for the real message, and a test asserts the second
attempt never happens.
Beside the stubs there is now a **real** reproduction,
`test_against_a_real_machine_with_no_trust_store`, opt-in behind
`BEHAVISION_NETWORK_TESTS=1` because it reaches github.com. Python's default
context honours `SSL_CERT_FILE`, so pointing it at an empty file gives a
default context that trusts nobody — the python.org condition exactly — while
certifi is loaded by path and is unaffected.
It skips rather than passes where it cannot reproduce that, and the difference
is measured rather than assumed:
```
macOS Command Line Tools LibreSSL 2.8.3 128 CAs with SSL_CERT_FILE empty
(reads the system keychain)
python.org / pyenv build OpenSSL 3.5.8 0 CAs -> reproduces it
```
Checked for teeth by putting the shipped `except` back: both the corrected stub
test and the live one fail, and pass again when it is restored.
## "The cameras won't connect from my mobile internet" — and why that is not a bug
Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics, and the product was not giving it.
A camera lives on the shop's LAN behind a router. `192.168.1.121` means
*"something on the network I am attached to"* and nothing more — from mobile
data, a hotel, or head office it resolves to nobody, or to a completely
different device that happens to hold that number. There is no route in from
the internet and **there must not be**: an RTSP camera reachable from outside
is how a shop's cameras end up being watched by strangers.
This is the reason the product is split the way it is, and it is worth stating
plainly next to the split itself: the shop PC is the only machine on the
camera's LAN, so it does the connecting, and every other surface reaches it
**outbound** — the agent's pull, the arrivals feed, and `LiveHub`'s frame relay,
which is what makes **Watch live** work from anywhere while nothing connects in.
What was wrong is the message. `cannot reach 192.168.1.121:554 - Operation
timed out` reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. `_wrong_network_hint` now names the
cause, and it distinguishes two states that need opposite actions — the same
rule as `artifact` against `no_faces`, and `stale` against `not_connecting`:
| | |
|---|---|
| this computer **is** on that network | check the camera is powered on and that the address is right |
| this computer is **somewhere else** | the computer is in the wrong place; no setting here fixes it, recognition has to run on a machine in the shop |
- **The local address comes from a `connect`ed UDP socket** that sends nothing.
It only fixes a route so the kernel will name the source address — no packet
leaves, and it needs no dependency in an engine that already ships 200 MB of
models.
- **The LAN ranges are spelled out, not `is_private`.** That property is
broader than "an address on somebody's LAN": it also covers carrier-grade NAT
and the documentation networks (192.0.2, 198.51.100, 203.0.113), and telling
somebody who typed one of those that it is "on the shop's own network" is a
confident wrong answer in the place people look first. Found by a test using
`203.0.113.9` as its example of a *public* address, which `is_private` calls
private.
- **A DNS name gets no hint at all.** Nothing can be concluded about
`camera.local` from the string, and guessing is the failure mode this whole
message exists to fix.
- **With no network at all it still names the cause** and drops the comparison,
rather than claiming to know which network this machine is on.
## A tunnel for the demo: when it is the right answer, and when it is not
Asked after the mobile-internet question: *"what if we created a secure tunnel
or vpn, then the cameras in our office can be shown in the demo version too?"*
Three different jobs get confused here, and only one of them needs a tunnel.
**Seeing the estate from anywhere already works and needs nothing.** Viewer
mode plus `LiveHub` is exactly this: the shop PC pushes frames outbound and any
signed-in app sees them, from mobile data, a hotel or a customer's office. A
VPN would add an installation, a credential and a moving part to something that
already works with none of them.
**Demonstrating recognition is better done on the demo machine's own camera.**
`webcam: 0` picks a capture index instead of building an RTSP URL, and the
engine has supported it since the first version — `CameraStore` round-trips it,
`source()` returns the index, `safe_url()` reports `webcam:0`, and
`POST /api/cameras {"id":"laptop","webcam":0}` has always worked. **No screen
offered it**, which is the same gap this file already records for the customer
record and for per-camera tuning: the API could, the UI could not reach it.
It is now an option in the make picker, and it is the strongest demo available:
real faces, in the room, instant, depending on no network at all. A demo
pointed at a camera in another building depends on two internet connections and
a tunnel staying up while somebody is talking.
**The one case a tunnel genuinely answers** is running the engine on a remote
machine against the office's own cameras — LAN access to `192.168.1.121` from
somewhere that is not that LAN. A mesh VPN (Tailscale and similar) does that
honestly: a subnet router at the office, the client on the demo machine, and
the address works unchanged. Free at this scale, no port forwarding, no
exposed camera. Worth using for our *own* office when that is really the goal.
**It is still the wrong answer for the product**, and the reasons are not about
difficulty:
- It is per-site infrastructure — an account, a node, a key — on every shop PC
we ship, to replace something that already works over ordinary HTTPS.
- It gives head office **network-level access into a customer's LAN**. The
current design can read a camera's picture; a VPN can reach the customer's
till, their router and everything else on that network. That is a far larger
thing to be trusted with, and a far larger thing to have breached.
- The failure modes are worse and less legible: a tunnel that is down looks
like a camera that is down.
So: tunnel for our own office if we want the engine running remotely; nothing
at all for showing customers their shops; the laptop's own camera for showing
anybody what the product does.
### And the suite was quietly 32 tests smaller than it looked
Installing `httpx` to check the webcam path took the engine suite from **239
passed to 271**. `tests/test_api_cameras.py` and its siblings begin with
`pytest.importorskip("httpx")` so a bare checkout still runs — deliberate, and
it means the number at the bottom of a run is not the number of tests that
exist. `pip install -e .[dev]` is the opt-in.

View File

@@ -46,6 +46,61 @@ import (
// otherwise arrive as a syntax error deep inside a dependency.
const minMinor = 10
// maxMinor is a WHEEL-availability ceiling, not a language one, and it is the
// reason this constant exists at all.
//
// findPython used to take the newest interpreter it could find, with a floor
// and no ceiling - which is precisely backwards, because the newest Python is
// the one least likely to have binary wheels for anything. Measured on a
// second Mac: it chose Python 3.14, pip found no numpy wheel for cp314, fell
// back to building numpy from source, and produced
//
// ERROR: Unknown compiler(s): [['cc'], ['gcc'], ['clang'], ...]
//
// then, once the operator installed Xcode's command line tools to get past
// that, ten minutes of compiling ending in
//
// arm_neon.h:28:2: error: "<arm_neon.h> is intended only for ARM and
// AArch64 targets"
//
// Two screens of C compiler output, on a shop counter, for a version choice
// made silently by this program. Refusing in one line, before anything is
// downloaded, is the whole of the fix.
//
// Raise it when the dependency set has wheels for the next version. Today
// onnxruntime is the binding one (cp314 is its newest); numpy publishes
// further ahead, and opencv-python ships a stable-ABI wheel that covers
// everything. `pip download --only-binary=:all: -r requirements.txt` against
// a candidate interpreter is the check.
const maxMinor = 14
// The three answers a candidate interpreter can get. Three, not two: a
// version that is too new and one that is too old need opposite actions from
// the operator, and collapsing them tells somebody holding Python 3.14 to go
// and install a newer Python.
const (
verdictOK = "ok"
verdictTooOld = "old"
verdictTooNew = "new"
verdictUnknown = "unparseable"
)
func pythonVerdict(major, minor int, parsed bool) string {
switch {
case !parsed:
return verdictUnknown
case major != 3:
// Python 4 is not a version this has been tried against, and 2 is
// long gone. Neither is a thing to guess about.
return verdictTooNew
case minor < minMinor:
return verdictTooOld
case minor > maxMinor:
return verdictTooNew
}
return verdictOK
}
func main() {
if err := run(); err != nil {
fmt.Fprintf(os.Stderr, "\n Setup did not finish: %v\n\n", err)
@@ -267,7 +322,13 @@ func findPython() (string, string, error) {
// `python3` therefore told a Mac with Python 3.12 sitting on it to go and
// install Python - measured on this machine, which has 3.12 under
// ~/.local/opt and reported "Found, but too old: python3 3.9".
versions := []string{"3.14", "3.13", "3.12", "3.11", "3.10"}
// Newest first WITHIN the supported range. Newest overall is what broke
// this; a version nobody has built wheels for is not a better choice than
// one that works.
var versions []string
for v := maxMinor; v >= minMinor; v-- {
versions = append(versions, fmt.Sprintf("3.%d", v))
}
for _, v := range versions {
cands = append(cands, cand{"python" + v, nil})
}
@@ -290,7 +351,7 @@ func findPython() (string, string, error) {
}
}
var tried []string
var tried, tooNew []string
for _, c := range cands {
exe := c.exe
if filepath.IsAbs(exe) {
@@ -314,7 +375,18 @@ func findPython() (string, string, error) {
}
ver := strings.TrimSpace(string(out))
tried = append(tried, c.exe+" "+ver)
if major, minor, ok := parseVer(ver); ok && (major > 3 || (major == 3 && minor >= minMinor)) {
major, minor, parsed := parseVer(ver)
switch verdict := pythonVerdict(major, minor, parsed); verdict {
case verdictTooNew:
// Recorded separately: "too new" and "too old" need opposite
// actions, and a single "found, but unsuitable" list sends
// somebody to upgrade a Python that is already past the problem.
tooNew = append(tooNew, c.exe+" "+ver)
continue
case verdictTooOld, verdictUnknown:
continue
}
{
full := exe
if len(c.args) > 0 {
full = exe + " " + strings.Join(c.args, " ")
@@ -327,7 +399,30 @@ func findPython() (string, string, error) {
// python.exe to PATH" on a Windows installer page reads as software that
// does not know where it is running, which is exactly the moment somebody
// stops trusting the rest of what it says.
msg := "no Python 3.10 or newer was found on this computer.\n\n"
// Only a too-new Python is a different problem with a different fix, and
// saying "no Python was found" to somebody looking at Python 3.14 is the
// kind of message that makes people stop believing the next one.
if len(tooNew) > 0 && len(tried) == 0 {
// Built as a value and wrapped, not written as an fmt.Errorf literal:
// this is a paragraph shown to an operator, and a linter that wants
// error strings to be lower-case fragments is right about errors
// programs read and wrong about the ones people do.
tooNewMsg := fmt.Sprintf(
"this computer has %s, which is newer than Behavision supports.\n\n"+
" Some of the libraries the engine needs have no build for it\n"+
" yet, so installing would fail part-way through.\n\n"+
" Install Python 3.%d and run this again:\n"+
" macOS: brew install python@3.%d\n"+
" or https://www.python.org/downloads/macos/\n"+
" Windows: https://www.python.org/downloads/windows/\n\n"+
" Both versions can sit on the machine together; this picks\n"+
" the one it can use.",
strings.Join(tooNew, ", "), maxMinor, maxMinor)
return "", "", errors.New(tooNewMsg)
}
msg := fmt.Sprintf("no Python between 3.%d and 3.%d was found on this computer.\n\n",
minMinor, maxMinor)
if runtime.GOOS == "windows" {
msg += " Install it from https://www.python.org/downloads/windows/\n" +
" and tick \"Add python.exe to PATH\" on the first screen,\n" +
@@ -339,6 +434,9 @@ func findPython() (string, string, error) {
if len(tried) > 0 {
msg += "\n\n Found, but too old: " + strings.Join(tried, ", ")
}
if len(tooNew) > 0 {
msg += "\n\n Found, but too new: " + strings.Join(tooNew, ", ")
}
return "", "", errors.New(msg)
}
@@ -375,14 +473,52 @@ func venvPython(venv string) string {
// engine requires - inside a shared interpreter is how you break the other
// thing months later, silently.
func makeVenv(py, venv string) error {
// An existing environment is reused - but only if the Python inside it is
// one this build supports.
//
// It used to be reused unconditionally, and that would have made the
// version ceiling above look like it did not work. The machine this was
// all found on already had a runtime built by Python 3.14, from the run
// that failed: with the ceiling in place setup would choose a good
// interpreter, reach here, find the 3.14 environment, keep it, and die in
// the same clang error as before. A fix that is defeated by the wreckage
// of the bug it fixes is not one.
//
// Rebuilding costs a re-download of the libraries and nothing else. The
// models are in the state root, not in here, so they survive.
if _, err := os.Stat(venvPython(venv)); err == nil {
return nil // already built; pip below brings it up to date
ok, ver := venvUsable(venv)
if ok {
return nil // pip below brings it up to date
}
fmt.Printf(" [..] %-24s %s\n", "Rebuilding environment",
"the existing one uses "+ver+", which is not supported")
if err := os.RemoveAll(venv); err != nil {
return fmt.Errorf("removing the old environment at %s: %w", venv, err)
}
}
exe, args := splitLauncher(py)
args = append(args, "-m", "venv", venv)
return stream(exec.Command(exe, args...), "creating the virtual environment")
}
// venvUsable reports whether the interpreter already inside an environment is
// one this build supports, and what it is when it is not.
//
// An environment that cannot be asked counts as unusable: a half-created or
// truncated one answers nothing, and reusing it fails later in pip with an
// error about a package rather than about the environment.
func venvUsable(venv string) (bool, string) {
out, err := exec.Command(venvPython(venv), "-c",
"import sys;print('%d.%d'%sys.version_info[:2])").Output()
if err != nil {
return false, "an interpreter that will not run"
}
ver := strings.TrimSpace(string(out))
major, minor, parsed := parseVer(ver)
return pythonVerdict(major, minor, parsed) == verdictOK, "Python " + ver
}
func pipInstall(vpy, src string) error {
fmt.Println(" Installing the engine and its libraries. This downloads a few")
fmt.Println(" hundred megabytes and takes a while on a slow connection.")
@@ -403,7 +539,30 @@ func pipInstall(vpy, src string) error {
// 'behavision.egg-info': Read-only file system". Falling back to source
// copies it somewhere writable first, for the same reason.
if wheels, _ := filepath.Glob(filepath.Join(src, "behavision-*.whl")); len(wheels) > 0 {
return stream(exec.Command(vpy, "-m", "pip", "install", "--upgrade", wheels[0]),
// TWO calls, and the second is the one that matters.
//
// `--upgrade` alone is not an upgrade when the version has not moved:
// pip skips the wheel and says so, one line above this program
// printing "[ok] Engine and dependencies installed". The wheel version
// was a frozen literal for several releases, so every engine fix in
// them silently failed to reach any machine that had run setup once -
// while the Go binaries beside it, rebuilt every release, updated
// normally. Half the product current, half of it months old, and
// nothing saying which.
//
// release.sh stamps the tag into the version now, so the versions do
// differ. This does not rely on that: a rebuild at the same version is
// the ordinary case while developing, and "installed" has to mean the
// code in this folder either way.
if err := stream(exec.Command(vpy, "-m", "pip", "install", "--upgrade", wheels[0]),
"installing the engine"); err != nil {
return err
}
// --no-deps so this is our own package only: the call above has
// already settled the dependencies, and forcing those too would
// re-download ~300 MB of numpy, OpenCV and onnxruntime every run.
return stream(exec.Command(vpy, "-m", "pip", "install",
"--force-reinstall", "--no-deps", wheels[0]),
"installing the engine")
}
tmp, err := os.MkdirTemp("", "behavision-src-")

View File

@@ -0,0 +1,112 @@
package main
import (
"os"
"path/filepath"
"testing"
)
// The choice this program makes silently, and got wrong.
//
// findPython took the newest interpreter on the machine, with a floor and no
// ceiling - backwards, because the newest Python is the one least likely to
// have binary wheels. On a Mac holding Python 3.14 it chose 3.14, pip found
// no numpy wheel for cp314, fell back to a source build and produced two
// screens of clang errors ending in "<arm_neon.h> is intended only for ARM
// and AArch64 targets". The operator's machine was fine; the version was not.
func TestTooNewIsRefusedRatherThanCompiled(t *testing.T) {
if got := pythonVerdict(3, maxMinor+1, true); got != verdictTooNew {
t.Errorf("3.%d = %q, want %q - picking it means a source build",
maxMinor+1, got, verdictTooNew)
}
if got := pythonVerdict(3, maxMinor, true); got != verdictOK {
t.Errorf("3.%d = %q, want %q - the ceiling is inclusive", maxMinor, got, verdictOK)
}
}
// Too old and too new must stay different answers. Telling somebody holding
// Python 3.14 that no Python was found, or that theirs is too old, sends them
// to install a newer one - which is the direction that already failed.
func TestOldAndNewAreDifferentAnswers(t *testing.T) {
old := pythonVerdict(3, minMinor-1, true)
fresh := pythonVerdict(3, maxMinor+1, true)
if old == fresh {
t.Fatalf("3.%d and 3.%d both reported %q", minMinor-1, maxMinor+1, old)
}
if old != verdictTooOld {
t.Errorf("3.%d = %q, want %q", minMinor-1, old, verdictTooOld)
}
}
// Every version in the range is accepted, so the window this program claims
// to support is the one it actually uses.
func TestTheWholeSupportedRangeIsAccepted(t *testing.T) {
for m := minMinor; m <= maxMinor; m++ {
if got := pythonVerdict(3, m, true); got != verdictOK {
t.Errorf("3.%d = %q, want %q", m, got, verdictOK)
}
}
if minMinor > maxMinor {
t.Fatal("the supported range is empty; nothing would ever be chosen")
}
}
// A major version nobody has tested against is not something to guess at, and
// an unreadable version string is not a working interpreter.
func TestUnknownVersionsAreNotAccepted(t *testing.T) {
for _, c := range []struct {
name string
major, minor int
parsed bool
}{
{"python 4", 4, 0, true},
{"python 2", 2, 7, true},
{"unparseable", 0, 0, false},
} {
if got := pythonVerdict(c.major, c.minor, c.parsed); got == verdictOK {
t.Errorf("%s was accepted", c.name)
}
}
}
// An environment already on disk is reused, and that is right until the Python
// inside it is one this build cannot use.
//
// It was reused unconditionally, which would have defeated the ceiling above
// on the exact machine that found the bug: that Mac already had a runtime
// built by Python 3.14, left behind by the run that failed. Setup would pick a
// good interpreter, find the 3.14 environment, keep it, and die in the same
// clang error as before - a fix defeated by the wreckage of the bug it fixes.
//
// Real environments, not a fake: the thing under test is what an interpreter
// on disk reports about itself.
func TestAnUnsupportedEnvironmentIsNotReused(t *testing.T) {
py, _, err := findPython()
if err != nil {
t.Skipf("no supported Python on this machine: %v", err)
}
venv := filepath.Join(t.TempDir(), "runtime")
if err := makeVenv(py, venv); err != nil {
t.Fatalf("makeVenv: %v", err)
}
if ok, ver := venvUsable(venv); !ok {
t.Fatalf("an environment built from the interpreter setup just chose "+
"reported itself unusable (%s)", ver)
}
// The two states that must not be confused with a working one.
empty := filepath.Join(t.TempDir(), "gone")
if ok, _ := venvUsable(empty); ok {
t.Error("a missing environment was reported usable")
}
broken := filepath.Join(t.TempDir(), "broken")
if err := os.MkdirAll(filepath.Dir(venvPython(broken)), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(venvPython(broken), []byte("not an interpreter"), 0o755); err != nil {
t.Fatal(err)
}
if ok, ver := venvUsable(broken); ok {
t.Errorf("a half-created environment was reported usable (%s)", ver)
}
}

View File

@@ -325,7 +325,7 @@ func cmdRun() error {
cfg.ClientID, cfg.SiteID, cfg.BrokerURL)
client, err := mqtt.NewClient(mqtt.ClientOptions{
BrokerURL: cfg.BrokerURL,
ClientID: "behavision-" + cfg.ClientID + "-" + cfg.SiteID,
ClientID: cfg.MQTTClientID(),
Username: cfg.BrokerUsername, Password: cfg.BrokerPassword,
CAFile: cfg.BrokerCAFile, Log: logger,
})

View File

@@ -30,12 +30,12 @@ import (
// argument, and it is why the wanted-check comes first and the push stops the
// moment the server says the last viewer has gone.
type Live struct {
Engine *EngineClient
Cloud *CloudClient
Log *log.Logger
FPS float64
Width int
Quality int
Engine *EngineClient
Cloud *CloudClient
Log *log.Logger
FPS float64
Width int
Quality int
}
// Defaults, measured against the office camera rather than guessed.

View File

@@ -8,7 +8,9 @@
package config
import (
"crypto/rand"
"encoding/base64"
"encoding/hex"
"encoding/json"
"fmt"
"os"
@@ -83,6 +85,10 @@ type Config struct {
// Queue.
SpoolMax int `json:"spool_max"`
// InstallID distinguishes THIS installation from every other one claimed
// to the same site. See MQTTClientID.
InstallID string `json:"install_id,omitempty"`
path string
}
@@ -124,6 +130,14 @@ func Load(path string) (Config, error) {
return cfg, fmt.Errorf("config %s: %w", path, err)
}
cfg.path = path
// Minted on first load and written back, so an installation that predates
// this field gets one without anybody doing anything. Best effort: a
// read-only config still yields a working id for this run, it is simply
// not the same one next time.
if cfg.InstallID == "" {
cfg.InstallID = newInstallID()
_ = cfg.Save(path)
}
for _, field := range []*string{&cfg.BrokerPassword, &cfg.APIPassword,
&cfg.SessionToken, &cfg.SessionRefresh, &cfg.AgentToken} {
plain, err := reveal(*field)
@@ -210,3 +224,41 @@ func reveal(stored string) (string, error) {
}
return string(plain), nil
}
// MQTTClientID names this INSTALLATION, not this site.
//
// It was `behavision-<client>-<site>`, which is the same string on every
// computer claimed to one shop. MQTT requires client ids to be unique and a
// broker enforces it by disconnecting the older session when a new one
// arrives with the same id - so two machines on one site take turns kicking
// each other off, forever. Measured on a second Mac claimed to a live shop:
//
// broker connected / broker connection lost: EOF / broker connected / ...
//
// The damage is not confined to the new machine. The shop's own till is the
// other half of that loop, so somebody signing in on a laptop to look at the
// product stops the shop delivering visits - and nothing at either end says
// why, because from each side it reads as an unstable network.
//
// The site stays in the id because it is what a broker log is read by, and
// the random half is short for the same reason. `CleanSession(true)` means
// there is no session state for a changed id to strand.
func (c Config) MQTTClientID() string {
id := c.InstallID
if id == "" {
// A config that could not be written still has to produce a UNIQUE
// id, or this falls straight back into the collision it exists to
// prevent. Per-run is the right failure: the connection works and the
// only cost is a new name in the broker's log after a restart.
id = newInstallID()
}
return "behavision-" + c.ClientID + "-" + c.SiteID + "-" + id
}
func newInstallID() string {
b := make([]byte, 4)
if _, err := rand.Read(b); err != nil {
return "x"
}
return hex.EncodeToString(b)
}

View File

@@ -0,0 +1,88 @@
package config
import (
"path/filepath"
"strings"
"testing"
)
// The bug this exists to prevent, measured on a second Mac claimed to a live
// shop: MQTT requires client ids to be unique, and a broker enforces it by
// disconnecting the older session when a new one arrives with the same id. The
// id was `behavision-<client>-<site>` - identical on every computer claimed to
// one shop - so the two took turns kicking each other off:
//
// broker connected / broker connection lost: EOF / broker connected / ...
//
// The damage is not confined to the new machine. The shop's own till is the
// other half of that loop, so somebody signing in on a laptop to look at the
// product stops the shop delivering visits.
func TestTwoInstallsOnOneSiteGetDifferentClientIDs(t *testing.T) {
dir := t.TempDir()
one := writeClaimed(t, filepath.Join(dir, "a.json"))
two := writeClaimed(t, filepath.Join(dir, "b.json"))
if one.MQTTClientID() == two.MQTTClientID() {
t.Fatalf("both installs answered to %q; the broker will disconnect one "+
"whenever the other connects", one.MQTTClientID())
}
}
// And the same install keeps its name across restarts, or a broker log is a
// list of strangers and nobody can tell one till from a stream of new ones.
func TestOneInstallKeepsItsClientIDAcrossRestarts(t *testing.T) {
path := filepath.Join(t.TempDir(), "agent.json")
first := writeClaimed(t, path)
again, err := Load(path)
if err != nil {
t.Fatalf("reload: %v", err)
}
if got, want := again.MQTTClientID(), first.MQTTClientID(); got != want {
t.Errorf("after a restart the id was %q, want %q", got, want)
}
}
// The site stays in the id: it is what somebody reading a broker log is
// reading FOR, and an opaque random string would make every connection
// anonymous.
func TestTheClientIDStillNamesTheShop(t *testing.T) {
c := Config{ClientID: "tenext-retail", SiteID: "chennai", InstallID: "abcd1234"}
id := c.MQTTClientID()
for _, want := range []string{"tenext-retail", "chennai", "abcd1234"} {
if !strings.Contains(id, want) {
t.Errorf("client id %q does not contain %q", id, want)
}
}
}
// A config that could not be written still has to produce a UNIQUE id, or a
// read-only install falls straight back into the collision. Per-run is the
// right failure: the connection works, and the only cost is a new name in the
// broker's log after a restart.
func TestAnUnsavedConfigStillGetsAUniqueID(t *testing.T) {
a := Config{ClientID: "c", SiteID: "s"}
b := Config{ClientID: "c", SiteID: "s"}
if a.MQTTClientID() == b.MQTTClientID() {
t.Fatal("two configs with no install id produced the same client id")
}
}
func writeClaimed(t *testing.T, path string) Config {
t.Helper()
cfg := Defaults()
cfg.ClientID, cfg.SiteID = "tenext-retail", "chennai"
if err := cfg.Save(path); err != nil {
t.Fatalf("save: %v", err)
}
// Loading is what mints the id, so an installation that predates the
// field gets one without anybody doing anything.
got, err := Load(path)
if err != nil {
t.Fatalf("load: %v", err)
}
if got.InstallID == "" {
t.Fatal("loading a config without an install id did not mint one")
}
return got
}

View File

@@ -4,4 +4,15 @@ Pipeline: capture -> detect (YuNet) -> track (IoU) -> align + encode
(ArcFace ONNX) -> match / auto-enroll (FAISS + SQLite) -> events + API.
"""
__version__ = "1.0.0"
# Stamped by release.sh from the git tag, into a copy of this line, so a
# release always produces a wheel nobody has installed before. It was a frozen
# "1.0.0" here and a frozen "1.1.0" in pyproject.toml - two literals that
# disagreed with each other and tracked nothing - which is how every engine fix
# for several releases silently failed to reach a machine that had run setup
# once. pip skips a wheel whose version is already installed and says so, one
# line above setup printing "[ok] Engine and dependencies installed".
#
# In a checkout it stays obviously a checkout: "0.0.0+dev" on /api/health is
# the truth about a developer's machine, and a plausible-looking number there
# would be worse than none.
__version__ = "0.0.0+dev"

View File

@@ -17,6 +17,7 @@ from fastapi.responses import HTMLResponse, Response, StreamingResponse
from fastapi.security import HTTPBasic, HTTPBasicCredentials
from pydantic import BaseModel, ValidationError
from . import __version__
from .config import ApiSection, CameraConfig, CameraTuning
from .commission import CommissionRun
from .events import Event
@@ -142,7 +143,7 @@ def _reencode(jpeg: bytes, width: int, quality: int) -> "bytes | None":
def create_app(engine: Engine) -> FastAPI:
app = FastAPI(title="Behavision", version="1.0.0",
app = FastAPI(title="Behavision", version=__version__,
dependencies=_auth_dependencies(engine.cfg.api))
def worker_or_404(camera_id: str):
@@ -172,6 +173,10 @@ def create_app(engine: Engine) -> FastAPI:
# are up, and the shop recognises nobody it already knows.
stranded = engine.gallery.health["stranded"]
return {"status": "ok" if engine.started_at else "starting",
# Which build is actually running. Without it there was no way
# to tell a shop PC three releases behind from a current one -
# which is precisely how a silent install failure survived.
"version": __version__,
"recognition_model": engine.encoder.model_name,
"gallery_unreadable_embeddings": stranded,
# "where is my database" must be answerable from the API: the

View File

@@ -68,6 +68,78 @@ os.environ.setdefault(
STALL_AFTER_S = 10.0
def _local_ipv4() -> str:
"""This machine's address on the interface holding the default route.
A UDP socket is `connect`ed and nothing is sent - it only fixes a route so
the kernel will name the source address. No packet leaves, and it needs no
dependency, which matters in an engine that already ships 200 MB of models.
"""
import socket
try:
with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as s:
s.settimeout(0.5)
s.connect(("8.8.8.8", 80))
return s.getsockname()[0]
except OSError:
return ""
def _wrong_network_hint(host: str) -> str:
"""Why a private camera address is unreachable, when that is the reason.
A camera lives on the shop's LAN behind a router, and 192.168.x.x means
"something on the network I am attached to" - nothing more. From mobile
data, a hotel, or head office it either resolves to nobody or to a
completely different device that happens to hold that number. There is no
route in from the internet and there must not be: an RTSP camera reachable
from outside is how a shop's cameras end up being watched by strangers.
Without this the answer was "cannot reach 192.168.1.121:554 - Operation
timed out", which reads as a broken camera and sends somebody to re-type
an address and a password that were always correct. Asked directly by the
owner, about his own cameras, from his phone's connection.
Two states, two different actions, so they must not share a sentence: on
the same network the camera or its address is the problem; on a different
one the COMPUTER is in the wrong place and no setting will fix it.
"""
import ipaddress
try:
addr = ipaddress.ip_address(host)
except ValueError:
return "" # a DNS name; nothing can be concluded from the string
# The RFC1918 blocks and link-local, spelled out rather than `is_private`.
# That property is broader than "an address on somebody's LAN": it also
# covers the carrier-grade NAT range and the documentation networks
# (192.0.2, 198.51.100, 203.0.113), and telling somebody who typed one of
# those that it is "on the shop's own network" would be a confident wrong
# answer in the place people look first. Found by a test using 203.0.113.9
# as an example of a PUBLIC address, which `is_private` calls private.
lan = any(addr in ipaddress.ip_network(n) for n in
("10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "169.254.0.0/16")
if addr.version == 4)
if not lan:
return ""
mine = _local_ipv4()
if not mine:
return (" - that is a private address, reachable only from inside "
"the network the camera is on")
try:
same = ipaddress.ip_network(f"{mine}/24", strict=False).supernet_of(
ipaddress.ip_network(f"{host}/24", strict=False))
except (ValueError, TypeError):
same = False
if same:
return (f" - this computer is on that network ({mine}), so check the "
f"camera is powered on and that {host} is its address")
return (f" - this computer is on {mine}, not the camera's network. A "
f"private address like {host} is only reachable from inside the "
f"shop's own network, never over the internet or mobile data, so "
f"recognition has to run on a computer in the shop")
def _tcp_reachable(source: "str | int", timeout: float
) -> "tuple[bool, str]":
"""Cheap pre-flight for an rtsp:// URL. Non-URL sources pass through."""
@@ -85,10 +157,11 @@ def _tcp_reachable(source: "str | int", timeout: float
return True, ""
except socket.timeout:
return False, (f"no response from {parsed.hostname}:{port} within "
f"{timeout:.0f}s - check the IP address and that the "
f"camera is on the same network")
f"{timeout:.0f}s{_wrong_network_hint(parsed.hostname)}")
except OSError as exc:
return False, f"cannot reach {parsed.hostname}:{port} - {exc.strerror or exc}"
return False, (f"cannot reach {parsed.hostname}:{port} - "
f"{exc.strerror or exc}"
f"{_wrong_network_hint(parsed.hostname)}")
def _fourcc(cap) -> str:

View File

@@ -4,6 +4,8 @@ from __future__ import annotations
import logging
import shutil
import ssl
import urllib.error
import urllib.request
from pathlib import Path
@@ -35,6 +37,81 @@ _COPY_MAP = {
}
def _https_context() -> "ssl.SSLContext | None":
"""The CA store to trust, or None to use whatever Python defaults to.
Returning None first is deliberate. On Windows and on a Homebrew or
system Python, the default context reads the machine's own certificate
store - which is what makes a corporate proxy with its own root CA work.
Replacing that with certifi's bundle unconditionally would break every
site that has one, in order to fix a different platform.
The platform this fixes is a python.org macOS build. It ships its own
OpenSSL with NO trust store, and populates one only when somebody
double-clicks `Install Certificates.command` in the Python folder -
which nobody installing face-recognition software has any reason to know
about. Every HTTPS request from that interpreter fails with:
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
certificate verify failed: unable to get local issuer certificate
Measured on a colleague's Mac: the engine installed perfectly and then
could not download a 230 KB model file, ending setup in forty lines of
traceback about `_ssl.c`.
"""
try:
import certifi
except ImportError: # pragma: no cover - certifi ships with requests
return None
return ssl.create_default_context(cafile=certifi.where())
def _is_cert_failure(err: BaseException) -> bool:
"""Is this a certificate-verification failure, however it is wrapped?
`urllib` does NOT let `ssl.SSLCertVerificationError` out. It catches it and
re-raises `urllib.error.URLError(err)`, carrying the original on `.reason`
- so `except ssl.SSLCertVerificationError` around `urlopen` matches
nothing, ever.
That is not a subtlety this file gets to record academically: the first
version of the fallback below was written exactly that way, shipped, and
failed on the machine it was written for with the very traceback it was
meant to prevent. The unit test passed throughout, because the fake it
used raised the bare SSL error - a shape real urllib never produces. A
stub that agrees with the author is worse than no test, and the test now
raises what urllib raises.
"""
reason = getattr(err, "reason", None)
return isinstance(err, ssl.SSLCertVerificationError) or \
isinstance(reason, ssl.SSLCertVerificationError)
def _urlopen(url: str, timeout: float = 60.0):
"""Open a URL, falling back to certifi's CA bundle on a verify failure.
Default first, certifi second, so the fix is additive: a machine whose
own store works keeps using it, and one with no store at all gets a
bundle rather than a traceback. certifi is already here - `requests` is a
hard dependency and brings it.
"""
try:
return urllib.request.urlopen(url, timeout=timeout)
except (urllib.error.URLError, ssl.SSLCertVerificationError) as err:
# Only a certificate problem is worth a second attempt. "No route to
# host" and "connection refused" arrive as URLError too, and retrying
# those with a different CA list changes nothing except how long the
# operator waits for the real message.
if not _is_cert_failure(err):
raise
ctx = _https_context()
if ctx is None:
raise
log.info("the system certificate store could not verify %s; "
"using the bundled CA list", url.split("/")[2])
return urllib.request.urlopen(url, timeout=timeout, context=ctx)
def _fetch(url: str, dest: Path, label: str) -> None:
"""Download with progress on stdout the supervisor can read.
@@ -56,7 +133,22 @@ def _fetch(url: str, dest: Path, label: str) -> None:
last = pct
log.info("download: %s %d%%", label, pct)
urllib.request.urlretrieve(url, dest, hook)
# Streamed rather than urlretrieve, only because urlretrieve offers no way
# to pass an SSL context and the whole point here is choosing one. The
# `download: <label> <n>%` lines are a contract: the supervisor parses
# them (`progressRe`) to put first-run progress in the tray, and without
# them a shop PC shows a stopped engine for five minutes after install.
with _urlopen(url) as resp:
total = int(resp.headers.get("Content-Length") or 0)
blocks, block_size = 0, 64 * 1024
with open(dest, "wb") as out:
while True:
chunk = resp.read(block_size)
if not chunk:
break
out.write(chunk)
blocks += 1
hook(blocks, block_size, total)
log.info("download: %s 100%%", label)
@@ -91,7 +183,7 @@ def setup_models(models_dir: Path) -> "list[str]":
import io
import zipfile
with urllib.request.urlopen(BUFFALO_SC_URL) as resp:
with _urlopen(BUFFALO_SC_URL) as resp:
payload = io.BytesIO(resp.read())
with zipfile.ZipFile(payload) as zf, \
zf.open("w600k_mbf.onnx") as src, \

View File

@@ -44,6 +44,9 @@ type App struct {
broker *agentmqtt.Client
stopBridge func()
hookURL string
// The resolved engine command, so engineMissing() and the supervisor are
// never looking at two different paths.
engineExe string
// Relays camera feeds to the webview so the engine's credential never has
// to travel in an <img> src, which a Chromium webview would strip anyway.
proxy *streamProxy
@@ -82,6 +85,14 @@ func (a *App) startup(ctx context.Context) {
if err := a.proxy.start(a.local.Base, a.local.User, a.local.Password); err != nil {
log.Printf("camera relay unavailable, tiles will not load: %v", err)
}
// And the other direction: watching a camera in another building, through
// head office's relay. Enabled unconditionally rather than only when a
// session already exists, because signing in is a thing that happens
// while the app is open - and CameraLive refuses without a session
// anyway, so there is nothing to gate.
if err := a.proxy.watchRemote(a.cloud.CameraLive); err != nil {
log.Printf("remote camera view unavailable: %v", err)
}
// A saved session means a shop PC that rebooted overnight comes back
// working instead of waiting for someone to log in.
@@ -101,6 +112,9 @@ func (a *App) startup(ctx context.Context) {
if exe != "" && !filepath.IsAbs(exe) {
exe = filepath.Join(agentpaths.InstallRoot(), exe)
}
a.mu.Lock()
a.engineExe = exe
a.mu.Unlock()
logFile, _ := agentengine.LogFile(agentpaths.EngineLog())
a.sup = agentengine.New(agentengine.Options{
Command: func(c context.Context) *exec.Cmd {
@@ -146,13 +160,55 @@ func (a *App) startup(ctx context.Context) {
// not run yet, starting the supervisor would loop on a missing executable
// with nothing useful to say. The Start button still exists for the one
// case where somebody has deliberately stopped it.
if _, err := os.Stat(exe); err == nil {
if why := a.engineMissing(); why == "" {
a.sup.Start()
} else {
log.Printf("engine not installed yet (%s); run behavision-setup, then Start", exe)
log.Printf("%s (looked for %s)", why, exe)
}
}
// engineMissing says, in a sentence somebody can act on, why recognition
// cannot start here - or "" when it can.
//
// It exists because the answer was only ever given at startup, to a log file
// nobody on a shop counter opens. Pressing Start went straight to the
// supervisor, which reported what exec reported:
//
// engine failed to start: fork/exec /private/var/folders/c2/.../
// AppTranslocation/500A5354-.../d/Behavision.app/Contents/MacOS/engine/
// behavision: no such file or directory
//
// Every word of that is true and none of it says "run the setup tool". One
// function, consulted by the startup path, the Start button and the status
// panel, so the three cannot give three different accounts of one fact.
func (a *App) engineMissing() string {
a.mu.RLock()
exe := a.engineExe
a.mu.RUnlock()
if exe == "" {
return "The recognition engine is not set up on this computer yet."
}
if _, err := os.Stat(exe); err == nil {
return ""
}
msg := "The recognition engine is not installed on this computer yet. " +
"Run behavision-setup from the folder you unzipped, then press Start."
// macOS quarantines a downloaded app it cannot verify and runs it from a
// randomly named READ-ONLY copy - App Translocation. Every relative path
// then resolves inside that copy, which is why the engine folder appears
// to be missing from a bundle that plainly contains one, and why an
// install into it would not survive a restart. Detectable, unguessable,
// and fixed by one drag; saying nothing leaves somebody re-running a
// setup tool that cannot win.
if strings.Contains(exe, "/AppTranslocation/") {
msg = "macOS is running Behavision from a temporary read-only copy, " +
"because it was opened straight from Downloads. Move Behavision " +
"to your Applications folder and open it from there, then run " +
"behavision-setup."
}
return msg
}
// webhookURL is the loopback address the bridge is listening on, or empty
// before it has started.
func (a *App) webhookURL() string {
@@ -234,7 +290,7 @@ func (a *App) startPipeline(ctx context.Context) {
}
client, err := agentmqtt.NewClient(agentmqtt.ClientOptions{
BrokerURL: a.cfg.BrokerURL,
ClientID: "behavision-" + a.cfg.ClientID + "-" + a.cfg.SiteID,
ClientID: a.cfg.MQTTClientID(),
Username: a.cfg.BrokerUsername, Password: a.cfg.BrokerPassword,
CAFile: a.cfg.BrokerCAFile, Log: logger,
})
@@ -587,6 +643,11 @@ func (a *App) EngineStatus() EngineStatus {
if err != nil {
out.Error = err.Error()
}
// The supervisor's own error is an exec failure; this replaces it with
// the reason, which is the part that tells somebody what to do.
if why := a.engineMissing(); why != "" {
out.Error = why
}
ctx, cancel := context.WithTimeout(a.ctx, 4*time.Second)
defer cancel()
// A running process is not a working engine: on a memory-starved box the
@@ -601,6 +662,13 @@ func (a *App) EngineStatus() EngineStatus {
}
func (a *App) StartEngine() EngineStatus {
// Refused rather than attempted. Handing a missing path to the supervisor
// produces a retry loop and an exec error for a message.
if why := a.engineMissing(); why != "" {
st := a.EngineStatus()
st.Error = why
return st
}
if a.sup != nil {
a.sup.Start()
}
@@ -646,6 +714,7 @@ func (a *App) Cameras() ([]map[string]any, error) {
"id": c.ID, "camera_id": c.CameraID, "label": c.Label,
"site": c.Site, "enabled": c.Enabled,
"connected": c.Connected, "last_seen_at": c.LastSeenAt,
"state": c.State, "state_note": c.StateNote,
"snapshot": c.Snapshot, "snapshot_at": c.SnapshotAt,
// What the screen keys off to hide Edit, Test and Check: this
// camera is on a network this PC cannot reach.
@@ -719,6 +788,23 @@ func (a *App) StreamURL(cameraID string) string {
return fmt.Sprintf("http://%s/api/cameras/%s/stream.mjpeg", base, cameraID)
}
// RemoteStreamURL is the live view of a camera in another building.
//
// The picture comes from head office's relay - the shop PC pushes frames
// outbound because nothing can reach in - and this app re-emits them as MJPEG
// on its own loopback, so a tile is an ordinary <img> either way. A screen
// therefore never has to know which building it is looking at.
//
// Empty when the relay is not running, and the caller shows the last snapshot
// instead. There is no useful fallback URL: the head-office endpoint needs
// this session's bearer, which an <img> cannot send.
func (a *App) RemoteStreamURL(cameraID string) string {
if !a.cloud.LoggedIn() {
return ""
}
return a.proxy.urlFor(cameraID, "live.mjpeg")
}
// ------------------------------------------------------------------- live --
type LiveSnapshot struct {

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View File

@@ -4,8 +4,8 @@
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>Behavision</title>
<script type="module" crossorigin src="./assets/index-vAtlvw9l.js"></script>
<link rel="stylesheet" crossorigin href="./assets/index-xWw5ie4A.css">
<script type="module" crossorigin src="./assets/index-p8f6baZq.js"></script>
<link rel="stylesheet" crossorigin href="./assets/index-DOJ2bRrM.css">
</head>
<body>
<div id="root"></div>

View File

@@ -38,6 +38,9 @@ export const api = {
startPlacement: (id, seconds) => call('StartPlacementCheck', id, seconds),
placementResult: (id) => call('PlacementResult', id),
streamURL: (id) => call('StreamURL', id),
// The live view of a camera in another building, relayed through head
// office. Empty when nobody is signed in.
remoteStreamURL: (id) => call('RemoteStreamURL', id),
live: () => call('Live'),
pipelineStatus: () => call('PipelineStatus'),

View File

@@ -648,3 +648,18 @@ tr.click { cursor: pointer; } tr.click:hover td { background: var(--s2); }
}
.viewing b { color: var(--ink); font-weight: 600; }
.viewing svg { flex: none; margin-top: 2px; color: var(--accent); }
/* Watch live sits over the picture, opposite the connection pill. It is on
the tile rather than in the button row because it is about the picture, and
because the row it would otherwise join is hidden on a remote camera. */
.camview .btn.watch {
position: absolute;
right: 10px;
bottom: 10px;
background: rgba(0, 0, 0, .55);
border-color: rgba(255, 255, 255, .25);
color: #fff;
backdrop-filter: blur(6px);
}
.camview .btn.watch:hover { background: rgba(0, 0, 0, .72); }
.camview .btn.watch.on { background: var(--accent); border-color: var(--accent); color: #fff; }

View File

@@ -2,7 +2,7 @@ import { useEffect, useRef, useState } from 'react'
import { api, message } from '../bridge.js'
import { usePolled } from '../hooks.js'
import * as Icon from '../ui/icons.jsx'
import { MAKES, makeById } from '../../../../shared/cameraMakes.js'
import { MAKES, makeById, parseRtspUrl } from '../../../../shared/cameraMakes.js'
// The camera screen is a picture, not a settings table.
//
@@ -14,6 +14,18 @@ import { MAKES, makeById } from '../../../../shared/cameraMakes.js'
// signed off through.
const BLANK = { id: '', host: '', port: 554, path: '', username: '', password: '', max_width: 1280 }
// "This computer's own camera" - a webcam or a built-in FaceTime camera.
//
// The engine has supported it since the first version (`webcam: 0` picks a
// capture index instead of building an RTSP URL) and no screen has ever
// offered it: another case of the API being able to do something the UI
// could not reach. It matters most for the thing it was missing from, which
// is showing the product to somebody. A laptop's own camera gives real
// recognition, of real faces, in the room, depending on no network at all -
// where pointing a demo machine at a camera in another building depends on
// two internet connections and a tunnel staying up while you talk.
const WEBCAM = 'webcam'
export default function Cameras() {
const { data, error, reload } = usePolled(() => api.cameras(), 8000)
const [editing, setEditing] = useState(null)
@@ -27,6 +39,19 @@ export default function Cameras() {
// than one that is absent.
const remote = cams.some(c => c.remote)
const streams = useStreamURLs(remote ? [] : cams)
// ONE camera at a time, and that is a cost decision rather than a layout
// one. A remote view makes the shop computer upload frames for as long as
// somebody is watching, so a grid that all went live at once would put an
// estate's worth of cameras on the wire because somebody opened a page.
const [watching, setWatching] = useState(null)
const [watchURL, setWatchURL] = useState('')
useEffect(() => {
let alive = true
if (!watching) { setWatchURL(''); return }
api.remoteStreamURL(watching).then(u => { if (alive) setWatchURL(u || '') })
.catch(() => { if (alive) setWatchURL('') })
return () => { alive = false }
}, [watching])
async function remove(cam) {
if (!confirm(`Remove ${cam.id}? Recognition from it stops immediately.`)) return
@@ -49,8 +74,9 @@ export default function Cameras() {
{remote && <div className="viewing">
<b>Viewing your shops from here.</b> These cameras are wired to the shop
computers, so they are set up and checked there. The picture is each
camera's most recent frame, not live video.
computers, so they are set up and checked there. Each tile shows that
camera's most recent frame; <b>Watch live</b> asks the shop computer to
send video for as long as you are looking.
</div>}
{error && <div className="err"><Icon.Warning size={15} />{error}</div>}
@@ -66,7 +92,10 @@ export default function Cameras() {
</div>
: <div className="camgrid">
{cams.map(c => (
<CameraCard key={c.id} cam={c} stream={streams[c.id]}
<CameraCard key={c.id} cam={c}
stream={watching === c.id ? watchURL : streams[c.id]}
watching={watching === c.id}
onWatch={() => setWatching(watching === c.id ? null : c.id)}
onEdit={() => setEditing(c)} onCheck={() => setCheck(c.id)} onRemove={() => remove(c)} />
))}
</div>}
@@ -78,14 +107,23 @@ export default function Cameras() {
)
}
function CameraCard({ cam, stream, onEdit, onCheck, onRemove }) {
function CameraCard({ cam, stream, watching, onWatch, onEdit, onCheck, onRemove }) {
// Three states, not two, and the third is why `connected` is a pointer on
// the wire: null means no shop computer has reported on this camera yet,
// which reads as waiting rather than as a fault to go and investigate.
const conn = cam.connected === undefined || cam.connected === null
? (cam.remote ? { tone: 'idle', label: 'Waiting for the shop computer' }
: { tone: 'idle', label: 'Engine stopped' })
: cam.connected ? { tone: 'ok', label: 'Connected' } : { tone: 'bad', label: 'Not connecting' }
const conn = cam.remote
// Four states, decided once by the server. `stale` is the one that was
// missing: the shop computer reports nothing when it cannot reach its own
// engine, so its last report used to sit there reading Connected -
// measured at 34 minutes on the live estate.
? ({ connected: { tone: 'ok', label: 'Connected' },
not_connecting: { tone: 'bad', label: 'Not connecting' },
stale: { tone: 'warn', label: 'Not reporting' } }[cam.state]
|| { tone: 'idle', label: 'Waiting for the shop computer' })
: cam.connected === undefined || cam.connected === null
? { tone: 'idle', label: 'Engine stopped' }
: cam.connected ? { tone: 'ok', label: 'Connected' }
: { tone: 'bad', label: 'Not connecting' }
// The last placement verdict, so "proven" survives closing the sheet. Only
// `good` is a pass: marginal means half the visitors are silently discarded.
// Never asked for a remote camera: that answer lives on the shop computer,
@@ -105,6 +143,9 @@ function CameraCard({ cam, stream, onEdit, onCheck, onRemove }) {
? <img src={stream || shot} alt={cam.id} />
: <div className="placeholder"><Icon.NoCamera size={34} /></div>}
<span className={`pill ${conn.tone === 'idle' ? '' : conn.tone} over`}><i className={`dot ${conn.tone}`} />{conn.label}</span>
{cam.remote && <button className={`btn sm watch ${watching ? 'on' : ''}`} onClick={onWatch}>
<Icon.Play size={13} />{watching ? 'Stop watching' : 'Watch live'}
</button>}
</div>
<div className="cambody">
<div className="camtitle">
@@ -121,9 +162,13 @@ function CameraCard({ cam, stream, onEdit, onCheck, onRemove }) {
</div>
{cam.remote
? <div className="camproof">
<span className="note">{cam.snapshot?.available
? 'Last picture from the shop computer.'
: cam.snapshot?.reason || 'No picture yet from the shop computer.'}</span>
<span className="note">{cam.state_note
? cam.state_note
: watching
? 'Live from the shop computer. It uploads only while you watch.'
: cam.snapshot?.available
? 'Last picture from the shop computer. Watch live to see it now.'
: cam.snapshot?.reason || 'No picture yet from the shop computer.'}</span>
</div>
: <div className="camproof">
<span className={`tag ${proof.tone}`}>{proof.label}</span>
@@ -157,7 +202,9 @@ function useStreamURLs(cams) {
function CameraSheet({ cam, onClose, onSaved }) {
const isNew = !cam.id
const [f, setF] = useState({ ...BLANK, ...cam, password: '', path: cam.path || (isNew ? MAKES[0].path : '') })
const [make, setMake] = useState(isNew ? MAKES[0].id : 'manual')
const [make, setMake] = useState(isNew ? MAKES[0].id : (cam.webcam != null ? WEBCAM : 'manual'))
const [index, setIndex] = useState(cam.webcam != null ? String(cam.webcam) : '0')
const local = make === WEBCAM
const [test, setTest] = useState(null)
const [busy, setBusy] = useState(null)
const [error, setError] = useState(null)
@@ -166,6 +213,7 @@ function CameraSheet({ cam, onClose, onSaved }) {
// Only overwrite the path when the preset has one, so choosing "I know the
// path" does not wipe what the installer already typed.
function chooseMake(e) {
if (e.target.value === WEBCAM) { setMake(WEBCAM); setTest(null); return }
const m = makeById(e.target.value)
setMake(m.id)
setF(prev => ({ ...prev, path: m.path || prev.path }))
@@ -179,6 +227,14 @@ function CameraSheet({ cam, onClose, onSaved }) {
if (v === '' || v === null || v === undefined) continue
out[k] = (k === 'port' || k === 'max_width') ? Number(v) : v
}
if (local) {
// An address and a webcam index are alternatives, not extras: the
// engine's source() takes the webcam first, so leaving a half-typed
// host behind would make the saved camera describe two different
// things and only one of them would be used.
for (const k of ['host', 'path', 'username', 'password']) delete out[k]
out.webcam = Number(index) || 0
}
return out
}
@@ -197,6 +253,27 @@ function CameraSheet({ cam, onClose, onSaved }) {
const chosen = makeById(make)
// Paste the whole RTSP URL. It is how people actually hold this
// information - it is what the camera's own app shows and what an installer
// writes down - and splitting it into five fields by eye is exactly where a
// password containing `@` or `/` goes wrong.
const [pasted, setPasted] = useState('')
const [pasteError, setPasteError] = useState('')
function applyUrl(text) {
setPasted(text)
if (!text.trim()) { setPasteError(''); return }
const got = parseRtspUrl(text)
if (!got) { setPasteError('That does not look like an RTSP address.'); return }
setPasteError('')
setTest(null)
// Only what the URL actually carried: a URL with no credentials must not
// wipe a password the operator typed above it.
setF(prev => ({ ...prev, host: got.host, port: got.port, path: got.path,
...(got.username ? { username: got.username } : {}),
...(got.password ? { password: got.password } : {}) }))
setMake('manual')
}
// The camera is picked from a scan of the shop's network rather than typed.
// Nobody knows their camera's address; the sticker is under the camera and
// the menu is different in every make's app. The scan names ONVIF cameras
@@ -226,7 +303,7 @@ function CameraSheet({ cam, onClose, onSaved }) {
<p className="lead">Three things from the camera: its address, its make, and its password. Test before you save — a wrong address is the most common mistake.</p>
{error && <div className="err"><Icon.Warning size={15} />{error}</div>}
{isNew && (
{isNew && !local && (
<section className="formsection">
<h4>Find it</h4>
{scan === null && (
@@ -260,6 +337,20 @@ function CameraSheet({ cam, onClose, onSaved }) {
</section>
)}
{isNew && !local && (
<section className="formsection">
<h4>Paste its address</h4>
<label className="field"><span>RTSP address</span>
<input className="mono" value={pasted} onChange={e => applyUrl(e.target.value)}
placeholder="rtsp://admin:password@192.168.1.20:554/ch0_0.264"
autoComplete="off" name="rtsp-url" spellCheck="false" />
<em className="hint">{pasteError
? pasteError
: 'If the camera’s own app shows an RTSP address, paste it here and the fields below fill in. Otherwise leave this empty and fill them in yourself.'}</em>
</label>
</section>
)}
<section className="formsection">
<h4>The camera</h4>
{isNew && (
@@ -268,15 +359,20 @@ function CameraSheet({ cam, onClose, onSaved }) {
<em className="hint">Short, no spaces. It names this camera everywhere and cannot be changed later.</em>
</label>
)}
<div className="fieldrow">
<label className="field"><span>Address</span>
<input value={f.host} onChange={set('host')} placeholder="192.168.1.20" inputMode="decimal" />
<em className="hint">On a label on the camera, or in its own app under “network”.</em>
</label>
<label className="field narrow"><span>Port</span>
<input value={f.port} onChange={set('port')} inputMode="numeric" />
</label>
</div>
{local
? <label className="field narrow"><span>Camera number</span>
<input value={index} onChange={e => { setIndex(e.target.value); setTest(null) }} inputMode="numeric" />
<em className="hint">0 is the built-in camera. Try 1 if a second one is plugged in.</em>
</label>
: <div className="fieldrow">
<label className="field"><span>Address</span>
<input value={f.host} onChange={set('host')} placeholder="192.168.1.20" inputMode="decimal" />
<em className="hint">On a label on the camera, or in its own app under “network”.</em>
</label>
<label className="field narrow"><span>Port</span>
<input value={f.port} onChange={set('port')} inputMode="numeric" />
</label>
</div>}
</section>
<section className="formsection">
@@ -284,27 +380,34 @@ function CameraSheet({ cam, onClose, onSaved }) {
<label className="field"><span>Make of camera</span>
<select value={make} onChange={chooseMake}>
{MAKES.map(m => <option key={m.id} value={m.id}>{m.label}</option>)}
<option value={WEBCAM}>This computer’s own camera</option>
</select>
{chosen.note && <em className="hint">{chosen.note}</em>}
</label>
<label className="field"><span>Stream path</span>
<input className="mono" value={f.path} onChange={set('path')} placeholder="/Streaming/Channels/101" />
<em className="hint">Filled in from the make. Change it only if the camera’s own app says something else.</em>
{local
? <em className="hint">Recognition runs on this computer’s built-in or plugged-in camera. Nothing on the network is involved.</em>
: chosen.note && <em className="hint">{chosen.note}</em>}
</label>
{!local && (
<label className="field"><span>Stream path</span>
<input className="mono" value={f.path} onChange={set('path')} placeholder="/Streaming/Channels/101" />
<em className="hint">Filled in from the make. Change it only if the camera’s own app says something else.</em>
</label>
)}
</section>
<section className="formsection">
<h4>Sign-in to the camera</h4>
<div className="fieldrow">
<label className="field"><span>Username</span>
<input name="rtsp-account" autoComplete="off" value={f.username} onChange={set('username')} placeholder="admin" />
</label>
<label className="field"><span>Password</span>
<input type="password" name="rtsp-secret" autoComplete="new-password" value={f.password}
onChange={set('password')} placeholder={cam.has_password ? '(unchanged)' : ''} />
</label>
</div>
</section>
{!local && (
<section className="formsection">
<h4>Sign-in to the camera</h4>
<div className="fieldrow">
<label className="field"><span>Username</span>
<input name="rtsp-account" autoComplete="off" value={f.username} onChange={set('username')} placeholder="admin" />
</label>
<label className="field"><span>Password</span>
<input type="password" name="rtsp-secret" autoComplete="new-password" value={f.password}
onChange={set('password')} placeholder={cam.has_password ? '(unchanged)' : ''} />
</label>
</div>
</section>
)}
{test && (
test.ok

View File

@@ -14,6 +14,7 @@ import (
"errors"
"fmt"
"io"
"net"
"net/http"
"net/url"
"strings"
@@ -682,6 +683,12 @@ type RemoteCamera struct {
Enabled bool `json:"enabled"`
Connected *bool `json:"connected"`
LastSeenAt string `json:"last_seen_at"`
// State is the server's single answer - connected / not_connecting /
// waiting / stale - and the screen renders that rather than deciding
// again from Connected. Two places deciding one fact is how a shop came
// out labelled Working, in green, above "2 of 3 cameras not connecting".
State string `json:"state"`
StateNote string `json:"state_note"`
Snapshot Photo `json:"snapshot"`
SnapshotAt string `json:"snapshot_at"`
}
@@ -746,3 +753,72 @@ func (c *Client) resolveShot(ctx context.Context, camID, at string, p Photo) Pho
p.URL, p.Auth = uri, false
return p
}
// CameraLive opens head office's live relay for one camera and returns the
// live SSE response for the caller to read and close.
//
// A response rather than frames, because the consumer is the app's own
// loopback relay: it re-emits these frames as MJPEG so an <img> can show them,
// and buffering the stream through a channel here would only add a place for
// frames to queue. A stale frame is worthless - the only one worth having is
// the newest - which is the whole reason LiveHub drops rather than queues.
//
// There is no client timeout on this request. A live view is endless by
// design and any deadline would cut the picture off mid-shift; the context is
// what ends it, when the viewer navigates away.
func (c *Client) CameraLive(ctx context.Context, cameraID string) (*http.Response, error) {
resp, err := c.liveOnce(ctx, cameraID)
if errors.Is(err, errTokenExpired) {
if rerr := c.Refresh(ctx); rerr != nil {
return nil, rerr
}
resp, err = c.liveOnce(ctx, cameraID)
}
return resp, err
}
func (c *Client) liveOnce(ctx context.Context, cameraID string) (*http.Response, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet,
c.Base+"/api/cameras/"+url.PathEscape(cameraID)+"/live", nil)
if err != nil {
return nil, err
}
req.Header.Set("Accept", "text/event-stream")
c.mu.RLock()
tok := c.token
c.mu.RUnlock()
if tok == "" {
return nil, ErrUnauthorized
}
req.Header.Set("Authorization", "Bearer "+tok)
// c.http has a 30 s timeout, which covers the whole response and would
// therefore sever a working live view every thirty seconds - the same
// trap that made the server set WriteTimeout to zero for its own SSE
// endpoint. A dedicated client, with the dial bounded instead.
hc := &http.Client{Transport: &http.Transport{
DialContext: (&net.Dialer{Timeout: 10 * time.Second}).DialContext,
TLSHandshakeTimeout: 10 * time.Second,
}}
resp, err := hc.Do(req)
if err != nil {
return nil, fmt.Errorf("cannot reach %s: %w", c.Base, err)
}
if resp.StatusCode == http.StatusUnauthorized {
var e struct {
Error string `json:"error"`
}
body, _ := io.ReadAll(io.LimitReader(resp.Body, 8192))
resp.Body.Close()
_ = json.Unmarshal(body, &e)
if e.Error == "token_expired" {
return nil, errTokenExpired
}
return nil, ErrUnauthorized
}
if resp.StatusCode >= 400 {
resp.Body.Close()
return nil, fmt.Errorf("live view: %s", resp.Status)
}
return resp, nil
}

View File

@@ -21,6 +21,7 @@ package main
// session and the bytes are fetched and handed over as an object URL.
import (
"context"
"crypto/rand"
"crypto/subtle"
"encoding/hex"
@@ -64,6 +65,12 @@ type streamProxy struct {
target string // engine origin, e.g. http://127.0.0.1:8010
user string
pass string
// Opens head office's live relay for one camera. Set on a computer that
// is signed in, whether or not an engine runs here - which is the whole
// point: watching a camera in another building is precisely the case
// where there is no engine on this machine to ask.
live func(ctx context.Context, cameraID string) (*http.Response, error)
}
func newStreamProxy() *streamProxy { return &streamProxy{} }
@@ -71,18 +78,49 @@ func newStreamProxy() *streamProxy { return &streamProxy{} }
// start binds a loopback listener and begins relaying. Calling it again while
// running is a no-op, so a restarted engine cannot leave two listeners behind.
func (p *streamProxy) start(base, user, pass string) error {
p.mu.Lock()
defer p.mu.Unlock()
if p.srv != nil {
return nil
}
if !strings.HasPrefix(base, "http://") && !strings.HasPrefix(base, "https://") {
base = "http://" + base
}
if _, err := url.Parse(base); err != nil {
return fmt.Errorf("engine base %q: %w", base, err)
}
if err := p.bind(); err != nil {
return err
}
p.mu.Lock()
defer p.mu.Unlock()
p.target = strings.TrimRight(base, "/")
p.user, p.pass = user, pass
return nil
}
// watchRemote makes the relay able to serve head office's live view, and
// binds it if nothing else has.
//
// Separate from start() because the two are independent: a shop PC has both
// an engine and a session, an owner's laptop has only a session, and a PC
// still being set up has only an engine. Folding them together would mean a
// computer with no engine could not watch a camera at all - which is the one
// computer most likely to be trying to.
func (p *streamProxy) watchRemote(fn func(context.Context, string) (*http.Response, error)) error {
if err := p.bind(); err != nil {
return err
}
p.mu.Lock()
defer p.mu.Unlock()
p.live = fn
return nil
}
// bind starts the loopback listener once. Calling it again while running is a
// no-op, so neither a restarted engine nor a second sign-in can leave two
// listeners behind.
func (p *streamProxy) bind() error {
p.mu.Lock()
defer p.mu.Unlock()
if p.srv != nil {
return nil
}
// The engine's own credential exists precisely so that the live face feed
// is never served open - CLAUDE.md is explicit that an unauthenticated
@@ -105,8 +143,6 @@ func (p *streamProxy) start(base, user, pass string) error {
p.ln = ln
p.token = hex.EncodeToString(raw)
p.target = strings.TrimRight(base, "/")
p.user, p.pass = user, pass
// No client timeout: an MJPEG stream is endless by design and any deadline
// would cut the picture off mid-shift. The request context ends it when
// the webview navigates away or the tile is replaced.
@@ -130,7 +166,7 @@ func (p *streamProxy) start(base, user, pass string) error {
func (p *streamProxy) stop() {
p.mu.Lock()
srv, ln := p.srv, p.ln
p.srv, p.ln, p.token = nil, nil, ""
p.srv, p.ln, p.token, p.live = nil, nil, "", nil
p.mu.Unlock()
if srv != nil {
_ = srv.Close()
@@ -155,6 +191,7 @@ func (p *streamProxy) urlFor(cameraID, file string) string {
func (p *streamProxy) handle(w http.ResponseWriter, r *http.Request) {
p.mu.RLock()
token, target, user, pass, client := p.token, p.target, p.user, p.pass, p.client
liveFn := p.live
p.mu.RUnlock()
if token == "" || client == nil {
http.NotFound(w, r)
@@ -190,6 +227,17 @@ func (p *streamProxy) handle(w http.ResponseWriter, r *http.Request) {
// calls; it is here so that adding a still later is a change to a screen
// rather than a change to the one file where a mistake is a credentialed
// proxy onto the biometric API.
// Head office's relay, not the engine. The two are different machines and
// different credentials, so this returns rather than falling through.
if parts[3] == "live.mjpeg" {
if liveFn == nil {
http.Error(w, "not signed in to head office", http.StatusBadGateway)
return
}
p.relayRemote(w, r, cameraID, liveFn)
return
}
var enginePath string
switch parts[3] {
case "stream.mjpeg":

149
desktop/stream_remote.go Normal file
View File

@@ -0,0 +1,149 @@
package main
// Watching a camera in another building, from the app.
//
// The shop PC sits behind a router with no inbound route, so nothing here can
// pull its MJPEG stream - that stream is served on the shop PC's own loopback
// and always will be. Head office's LiveHub is the way round it: the agent
// asks outbound whether anyone is watching and pushes JPEG frames up for
// exactly as long as somebody is. The head-office web app already consumes
// that; this is the same feed, for the app.
//
// It arrives as base64 frames over SSE, which an <img> cannot render, so this
// re-emits them as multipart MJPEG - which an <img> renders natively, through
// the relay that already exists for the local engine. That is what keeps ONE
// code path in the screens: a tile points at a loopback URL and does not know
// or care which building the picture came from.
import (
"bufio"
"context"
"encoding/base64"
"fmt"
"net/http"
"strings"
"time"
)
// The boundary is ours to choose; it only has to be a string the JPEG bytes
// cannot contain, and a marker line never appears inside JPEG data.
const mjpegBoundary = "behavisionframe"
// A frame is base64, so ~1.33 bytes on the wire per byte of picture. The
// engine re-encodes to 640 px for the relay and those measure ~20 KB, so this
// is roughly a hundredfold headroom - large enough never to clip a real frame
// and small enough that a broken or hostile stream cannot grow this process's
// memory without bound.
const maxFrameLine = 8 << 20
func (p *streamProxy) relayRemote(w http.ResponseWriter, r *http.Request,
cameraID string, open func(context.Context, string) (*http.Response, error)) {
w.Header().Set("Content-Type", "multipart/x-mixed-replace; boundary="+mjpegBoundary)
w.Header().Set("Cache-Control", "no-store")
flusher, _ := w.(http.Flusher)
// Send the headers NOW, before any frame exists. Go writes them on the
// first body write, so without this the whole response - status line
// included - waits for the shop computer to start pushing, and a viewer
// whose camera is slow to answer sees the REQUEST time out rather than a
// stream that has not painted yet. Measured against production: 30
// seconds and not even a Content-Type.
if flusher != nil {
flusher.Flush()
}
// Reconnecting is normal, not an error. The server caps one push at five
// minutes so that a tab left open for a week cannot leave a shop
// uploading for a week - so a viewer who IS still there simply asks
// again. Doing it here rather than in the page is what lets the <img>
// survive the cap: it never sees the stream end.
sent := 0
for {
if r.Context().Err() != nil {
return
}
n, err := p.pumpRemote(w, flusher, r.Context(), cameraID, open)
sent += n
if r.Context().Err() != nil {
return
}
// Nothing was written and the attempt failed. Writing an error body
// now would be writing it into a multipart stream the <img> is
// already parsing, so the picture simply stays on whatever it last
// showed and the screen's own "not connecting" state is the report.
if err != nil && sent == 0 {
return
}
select {
case <-r.Context().Done():
return
case <-time.After(1500 * time.Millisecond):
}
}
}
// pumpRemote runs one SSE connection to exhaustion and returns how many
// frames it forwarded.
func (p *streamProxy) pumpRemote(w http.ResponseWriter, flusher http.Flusher,
ctx context.Context, cameraID string,
open func(context.Context, string) (*http.Response, error)) (int, error) {
resp, err := open(ctx, cameraID)
if err != nil {
return 0, err
}
defer resp.Body.Close()
sc := bufio.NewScanner(resp.Body)
sc.Buffer(make([]byte, 0, 64*1024), maxFrameLine)
var event, data string
frames := 0
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event: "):
event = strings.TrimSpace(line[7:])
case strings.HasPrefix(line, "data: "):
data = line[6:]
case line == "":
// End of one SSE event. `waiting` means head office has us
// registered and the shop PC has not started pushing yet - a real
// second or two while the agent is asked, and nothing to draw.
if event == "frame" && data != "" {
if err := writeMJPEGFrame(w, flusher, data); err != nil {
return frames, err // the webview went away
}
frames++
}
event, data = "", ""
}
}
return frames, sc.Err()
}
func writeMJPEGFrame(w http.ResponseWriter, flusher http.Flusher, b64 string) error {
jpg, err := base64.StdEncoding.DecodeString(b64)
if err != nil || len(jpg) == 0 {
// One malformed frame is not a reason to tear down a working view.
return nil
}
if _, err := fmt.Fprintf(w,
"--%s\r\nContent-Type: image/jpeg\r\nContent-Length: %d\r\n\r\n",
mjpegBoundary, len(jpg)); err != nil {
return err
}
if _, err := w.Write(jpg); err != nil {
return err
}
if _, err := w.Write([]byte("\r\n")); err != nil {
return err
}
// Flushed per frame. Anything held waiting for a full buffer is a tile
// that stays blank, which is indistinguishable from the view not working.
if flusher != nil {
flusher.Flush()
}
return nil
}

View File

@@ -0,0 +1,103 @@
package main
import (
"bytes"
"context"
"net/http"
"os"
"testing"
"time"
"github.com/loyaly/behavision-desktop/internal/cloud"
)
// The whole chain against the real head office and a real shop computer:
//
// TEST_CLOUD_EMAIL=... TEST_CLOUD_PASSWORD=... \
// go test ./desktop/ -run RemoteLive -v
//
// Everything in stream_remote_test.go proves the relay against a fake that
// agrees with me. Only this proves the part that cannot be faked: that a shop
// computer behind a router with no inbound route actually pushes frames when
// asked, that they survive base64 and SSE, and that what comes out of the
// loopback relay is a multipart stream an <img> will paint.
//
// It also costs something to run, which is why it is opt-in: watching makes
// the shop computer upload for as long as the test reads.
func TestRemoteLiveFromProduction(t *testing.T) {
email, pass := os.Getenv("TEST_CLOUD_EMAIL"), os.Getenv("TEST_CLOUD_PASSWORD")
if email == "" || pass == "" {
t.Skip("set TEST_CLOUD_EMAIL and TEST_CLOUD_PASSWORD to run against production")
}
base := os.Getenv("TEST_CLOUD_URL")
if base == "" {
base = "https://mcp.loyaly.ai"
}
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
c := cloud.New(base)
if _, err := c.Login(ctx, email, pass); err != nil {
t.Fatalf("login: %v", err)
}
cams, err := c.RemoteCameras(ctx)
if err != nil {
t.Fatalf("cameras: %v", err)
}
t.Logf("%d cameras", len(cams))
target := os.Getenv("TEST_CLOUD_CAMERA")
for _, cam := range cams {
conn := "waiting"
if cam.Connected != nil {
conn = map[bool]string{true: "connected", false: "not connecting"}[*cam.Connected]
}
t.Logf(" %-10s %-16s %-15s snapshot=%v", cam.CameraID, cam.Site, conn, cam.Snapshot.Available)
if target == "" && cam.Connected != nil && *cam.Connected {
target = cam.CameraID
}
}
if target == "" {
t.Skip("no connected camera to watch")
}
p := newStreamProxy()
if err := p.watchRemote(c.CameraLive); err != nil {
t.Fatalf("watchRemote: %v", err)
}
defer p.stop()
rctx, rcancel := context.WithTimeout(ctx, 30*time.Second)
defer rcancel()
req, _ := http.NewRequestWithContext(rctx, http.MethodGet, p.urlFor(target, "live.mjpeg"), nil)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET relay: %v", err)
}
defer resp.Body.Close()
start := time.Now()
acc, buf, frames := make([]byte, 0, 1<<20), make([]byte, 32*1024), 0
for frames < 10 {
n, rerr := resp.Body.Read(buf)
acc = append(acc, buf[:n]...)
frames = bytes.Count(acc, []byte("--"+mjpegBoundary))
if rerr != nil {
break
}
}
el := time.Since(start)
t.Logf("watching %q: %d frames, %d bytes, %.1fs (%.1f fps, %.0f KB/s)",
target, frames, len(acc), el.Seconds(),
float64(frames)/el.Seconds(), float64(len(acc))/el.Seconds()/1024)
if frames < 3 {
t.Fatalf("got %d frames from a connected camera - the shop computer is "+
"not answering head office's request to push", frames)
}
// Bytes that are actually a picture, not a framing header that happens to
// be well formed. A JPEG begins FFD8.
if !bytes.Contains(acc, []byte{0xFF, 0xD8, 0xFF}) {
t.Error("no JPEG start marker anywhere in the stream")
}
}

View File

@@ -0,0 +1,209 @@
package main
import (
"bytes"
"context"
"encoding/base64"
"fmt"
"io"
"net/http"
"net/http/httptest"
"strings"
"sync/atomic"
"testing"
"time"
)
// jpg is a byte sequence that is not valid JPEG and does not need to be: what
// is under test is that the bytes arrive intact and framed, not that a decoder
// likes them.
var jpg = []byte{0xFF, 0xD8, 'h', 'e', 'l', 'l', 'o', 0xFF, 0xD9}
// sseServer answers head office's live endpoint with `pushes` frames and then
// ends the response, which is what the server's five-minute cap does.
func sseServer(t *testing.T, frames int, hits *int32) *httptest.Server {
t.Helper()
return httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
atomic.AddInt32(hits, 1)
w.Header().Set("Content-Type", "text/event-stream")
fl, _ := w.(http.Flusher)
// Registered, nothing being pushed yet. Nothing may be drawn for it.
fmt.Fprint(w, "event: waiting\ndata: \n\n")
if fl != nil {
fl.Flush()
}
for i := 0; i < frames; i++ {
fmt.Fprintf(w, "event: frame\ndata: %s\n\n",
base64.StdEncoding.EncodeToString(jpg))
if fl != nil {
fl.Flush()
}
}
}))
}
func openerFor(srv *httptest.Server) func(context.Context, string) (*http.Response, error) {
return func(ctx context.Context, cam string) (*http.Response, error) {
req, _ := http.NewRequestWithContext(ctx, http.MethodGet, srv.URL+"/live/"+cam, nil)
return http.DefaultClient.Do(req)
}
}
// The whole point: base64 frames over SSE are not something an <img> can show,
// and a multipart MJPEG stream is. Without this the app could only ever show a
// still, on exactly the computers that cannot reach the camera any other way.
func TestRemoteFramesReachTheWebviewAsMJPEG(t *testing.T) {
var hits int32
srv := sseServer(t, 3, &hits)
defer srv.Close()
p := newStreamProxy()
if err := p.watchRemote(openerFor(srv)); err != nil {
t.Fatalf("watchRemote: %v", err)
}
defer p.stop()
u := p.urlFor("cam2", "live.mjpeg")
if u == "" {
t.Fatal("no relay url; the proxy did not bind")
}
// The relay reconnects for as long as the viewer is there, so the read is
// bounded by us rather than by the stream ending - exactly as an <img>
// would behave.
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Second)
defer cancel()
req, _ := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET relay: %v", err)
}
defer resp.Body.Close()
if ct := resp.Header.Get("Content-Type"); !strings.HasPrefix(ct, "multipart/x-mixed-replace") {
t.Fatalf("Content-Type = %q, an <img> will not treat that as a stream", ct)
}
// Read the first three frames' worth and stop; the relay would otherwise
// go on reconnecting forever, which is the behaviour being relied on.
want := append([]byte(fmt.Sprintf("--%s\r\nContent-Type: image/jpeg\r\nContent-Length: %d\r\n\r\n",
mjpegBoundary, len(jpg))), jpg...)
got := make([]byte, 0, 4096)
buf := make([]byte, 512)
for len(got) < 3*len(want) {
n, rerr := resp.Body.Read(buf)
got = append(got, buf[:n]...)
if rerr != nil {
break
}
}
if n := bytes.Count(got, []byte("--"+mjpegBoundary)); n < 3 {
t.Fatalf("got %d frames in %d bytes, want at least 3", n, len(got))
}
if !bytes.Contains(got, want) {
t.Errorf("a frame was not framed as expected:\n%q", got[:min(len(got), 300)])
}
// `waiting` is a real state - head office has us registered and the shop
// computer has not started pushing - and there is nothing to draw for it.
// Emitting an empty part would blank a tile that already had a picture.
if bytes.Contains(got, []byte("Content-Length: 0")) {
t.Error("an empty frame was written for a waiting event")
}
}
// The server caps one push at five minutes so a tab left open for a week
// cannot leave a shop uploading for a week. Reconnecting is therefore a normal
// event, and doing it here rather than in the page is what lets the <img>
// survive the cap - it never sees the stream end.
func TestTheRelayReconnectsWhenHeadOfficeEndsAPush(t *testing.T) {
var hits int32
srv := sseServer(t, 1, &hits)
defer srv.Close()
p := newStreamProxy()
if err := p.watchRemote(openerFor(srv)); err != nil {
t.Fatalf("watchRemote: %v", err)
}
defer p.stop()
ctx, cancel := context.WithTimeout(context.Background(), 4*time.Second)
defer cancel()
req, _ := http.NewRequestWithContext(ctx, http.MethodGet, p.urlFor("cam2", "live.mjpeg"), nil)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET relay: %v", err)
}
defer resp.Body.Close()
// Two frames means two pushes, because each push carries exactly one.
seen, buf := 0, make([]byte, 256)
acc := make([]byte, 0, 2048)
for seen < 2 {
n, rerr := resp.Body.Read(buf)
acc = append(acc, buf[:n]...)
seen = bytes.Count(acc, []byte("--"+mjpegBoundary))
if rerr != nil {
break
}
}
if seen < 2 {
t.Fatalf("got %d frames across reconnects, want 2", seen)
}
if got := atomic.LoadInt32(&hits); got < 2 {
t.Errorf("head office was asked %d times, want at least 2", got)
}
}
// Signed out, the relay must not pretend. There is no fallback URL to offer
// either: the head-office endpoint needs this session's bearer, which an <img>
// cannot send - so a tile that silently failed would be the only alternative.
func TestTheRelayRefusesWhenNobodyIsSignedIn(t *testing.T) {
p := newStreamProxy()
if err := p.watchRemote(nil); err != nil {
t.Fatalf("watchRemote: %v", err)
}
defer p.stop()
resp, err := http.Get(p.urlFor("cam2", "live.mjpeg"))
if err != nil {
t.Fatalf("GET relay: %v", err)
}
defer resp.Body.Close()
io.Copy(io.Discard, resp.Body)
if resp.StatusCode != http.StatusBadGateway {
t.Errorf("status = %d, want 502", resp.StatusCode)
}
}
// The relay is credentialed - it is a path to a live view of a shop floor -
// and the token is the only thing standing between another local process and
// it. live.mjpeg must be behind exactly the same door as the engine routes.
func TestTheRemoteRouteIsBehindTheSameToken(t *testing.T) {
var hits int32
srv := sseServer(t, 1, &hits)
defer srv.Close()
p := newStreamProxy()
if err := p.watchRemote(openerFor(srv)); err != nil {
t.Fatalf("watchRemote: %v", err)
}
defer p.stop()
// The right shape, the wrong value.
parts := strings.Split(p.urlFor("cam2", "live.mjpeg"), "/")
parts[4] = strings.Repeat("0", len(parts[4]))
bad := strings.Join(parts, "/")
resp, err := http.Get(bad)
if err != nil {
t.Fatalf("GET relay: %v", err)
}
defer resp.Body.Close()
io.Copy(io.Discard, resp.Body)
if resp.StatusCode != http.StatusNotFound {
t.Errorf("status = %d, want 404 - and 404 rather than 403, because there is nothing here to tell an unwelcome caller they found the right door", resp.StatusCode)
}
if atomic.LoadInt32(&hits) != 0 {
t.Error("a request with the wrong token still made the shop computer upload")
}
}

View File

@@ -1,11 +1,12 @@
[project]
name = "behavision"
version = "1.1.0"
dynamic = ["version"]
description = "Production face recognition over RTSP"
requires-python = ">=3.10"
dependencies = [
"numpy>=1.26,<2.0",
"opencv-python>=4.8.1",
# See requirements.txt for why numpy is uncapped and opencv is not.
"numpy>=1.26,<3.0",
"opencv-python>=4.8.1,<5",
"onnxruntime>=1.16",
"fastapi>=0.110",
"uvicorn>=0.29",
@@ -21,7 +22,11 @@ dependencies = [
]
[project.optional-dependencies]
dev = ["pytest>=8.0"]
# httpx is test-only and never ships in the wheel. The HTTP tests begin with
# `pytest.importorskip("httpx")` so a bare checkout still runs - which is
# right, and has a cost worth knowing: without it the suite reports 239 passed
# and quietly SKIPS 32 API tests. `pip install -e .[dev]` is how to get them.
dev = ["pytest>=8.0", "httpx>=0.27"]
[tool.setuptools.packages.find]
include = ["behavision*"]
@@ -31,3 +36,9 @@ behavision = ["static/*"]
[tool.pytest.ini_options]
testpaths = ["tests"]
# One version for the package and the wheel. Two literals in two files is how
# they came to disagree (1.0.0 here, 1.1.0 there) and how neither tracked a
# release.
[tool.setuptools.dynamic]
version = {attr = "behavision.__version__"}

View File

@@ -50,7 +50,28 @@ step "2. Agent and setup tool"
step "3. Engine source and wheel"
# The wheel is built with the checkout's own interpreter; requires-python is a
# statement about the SHOP PC, which setup enforces when it finds Python there.
# Stamp the tag into the wheel version. Without this every release built
# behavision-1.1.0-py3-none-any.whl, and `pip install --upgrade` on a machine
# that already had 1.1.0 is a no-op - so an engine fix reached nobody who had
# ever run setup, while the Go binaries beside it updated normally. Nothing
# reported a version either, so there was no way to tell a shop PC three
# releases behind from a current one.
#
# v0.5.6-demo -> 0.5.6+demo, which is valid PEP 440: a local segment takes
# alphanumerics and dots, never hyphens.
TMPVER=$(mktemp)
PEP440=$(printf '%s' "${TAG#v}" | sed 's/-/+/; s/[^0-9A-Za-z.+]/./g')
trap 'git checkout -- behavision/__init__.py 2>/dev/null || true; rm -f "$TMPVER"' EXIT
# A temp file rather than `sed -i`, whose argument handling differs between
# BSD and GNU - this script is run from a Mac today and that is not a reason
# to plant a portability trap in a release path.
sed "s/^__version__ = .*/__version__ = \"$PEP440\"/" behavision/__init__.py > "$TMPVER"
cat "$TMPVER" > behavision/__init__.py
.venv/bin/python -m pip wheel --no-deps --ignore-requires-python -q -w "$STAGE/engine-src" . 2>&1 | grep -v "DEPRECATION\|WARNING: Ignoring" || true
git checkout -- behavision/__init__.py
rm -f "$TMPVER"
trap - EXIT
echo " engine wheel version $PEP440"
ls "$STAGE"/engine-src/behavision-*.whl >/dev/null || { echo "wheel was not built" >&2; exit 1; }
cp pyproject.toml requirements.txt "$STAGE/engine-src/"
mkdir -p "$STAGE/engine-src/config" && cp config/default.yaml "$STAGE/engine-src/config/"

View File

@@ -1,5 +1,31 @@
numpy>=1.26,<2.0
opencv-python>=4.8.1
# numpy 2 is allowed, and that is what lets this install on a current Python.
# `<2.0` capped the resolver at numpy 1.26.4, whose newest wheel is cp312, so
# on a Mac with Python 3.14 pip fell back to BUILDING numpy from source and
# died in clang - two screens of C compiler output on a shop counter, for a
# version choice made silently by behavision-setup.
#
# Measured before changing it, nine runs of the detector guard per combination
# on one machine:
#
# numpy 1.26 / cv2 4.11 9 passed, 0 crashed
# numpy 2.0 / cv2 4.11 8 passed, 1 crashed
# numpy 1.26 / cv2 4.14 3 passed, 6 crashed
# numpy 2.0 / cv2 4.14 2 passed, 7 crashed
#
# numpy is not the variable there; OpenCV is. The crash was a test racing a
# shared cv2.FaceDetectorYN on purpose - undefined behaviour in C++, which 4.11
# usually turned into an exception and 4.14 usually turns into a segfault. The
# product never shares one (Engine._build_worker builds a detector per camera),
# and the test now runs that race out of process. With that fixed, the whole
# suite is 226 passed / 2 skipped on numpy 2.0.2, five runs out of five.
#
# opencv IS capped, and the two are not the same call. Everything above was
# measured on 4.x; OpenCV 5.0 is a major release this project has never run a
# real camera or an emotion model through, and an uncapped `>=4.8.1` means
# every NEW install silently gets it while every existing one keeps 4.11. Lift
# it after running a camera on 5.x, not before.
numpy>=1.26,<3.0
opencv-python>=4.8.1,<5
onnxruntime>=1.16
fastapi>=0.110
uvicorn>=0.29

View File

@@ -0,0 +1,96 @@
package api
import (
"testing"
"time"
)
func ptr(b bool) *bool { return &b }
// The state this was written for. Two cameras read "Connected", in green, on
// the live estate thirty-four minutes after the shop computer had stopped
// being able to see either of them - because the agent correctly reports
// nothing when it cannot reach the engine, and the last value it sent stays
// in the database looking current.
func TestAStaleReportIsNotAConnectedCamera(t *testing.T) {
now := time.Date(2026, 9, 30, 11, 13, 0, 0, time.UTC)
c := Camera{Connected: ptr(true),
LastSeenAt: now.Add(-34 * time.Minute).Format(time.RFC3339)}
c.CameraState(now)
if c.State != CameraStale {
t.Errorf("state = %q, want %q", c.State, CameraStale)
}
// Cleared, not merely overruled. A stale true left in place stays
// available to every client that reads the field directly, and leaves two
// fields on one object disagreeing.
if c.Connected != nil {
t.Errorf("connected = %v, want null - nobody currently knows", *c.Connected)
}
if c.StateNote == "" {
t.Error("a stale camera said nothing about what to do")
}
}
// One missed report is a dropped packet. Warning on it would put an alarm on a
// healthy estate every few minutes, and an indicator that cries wolf is one
// people learn to ignore.
func TestOneMissedReportIsStillConnected(t *testing.T) {
now := time.Now()
c := Camera{Connected: ptr(true),
LastSeenAt: now.Add(-90 * time.Second).Format(time.RFC3339)}
c.CameraState(now)
if c.State != CameraConnected {
t.Errorf("state = %q after 90s, want %q", c.State, CameraConnected)
}
if c.Connected == nil || !*c.Connected {
t.Error("a fresh report lost its connected flag")
}
}
// "Nobody has ever told us" and "nobody has told us lately" need different
// sentences: the first is a camera head office added a minute ago and the
// shop computer has not picked up, the second is a shop computer that has
// stopped. Sending an installer to the wrong one wastes a journey.
func TestNeverReportedIsNotTheSameAsStopped(t *testing.T) {
now := time.Now()
var never Camera
never.CameraState(now)
if never.State != CameraWaiting {
t.Errorf("state = %q, want %q", never.State, CameraWaiting)
}
stopped := Camera{LastSeenAt: now.Add(-time.Hour).Format(time.RFC3339)}
stopped.CameraState(now)
if stopped.State == never.State {
t.Fatal("a camera nobody has reported and one that stopped read the same")
}
if stopped.StateNote == never.StateNote {
t.Error("two states that need different actions gave the same advice")
}
}
// A camera the shop computer CAN see and cannot open is the one case where
// "check the cabling" is the right advice, and it must stay distinguishable
// from the three where it is not.
func TestAFreshFailureSaysCheckTheCamera(t *testing.T) {
now := time.Now()
c := Camera{Connected: ptr(false), LastSeenAt: now.Format(time.RFC3339)}
c.CameraState(now)
if c.State != CameraNotConnect {
t.Errorf("state = %q, want %q", c.State, CameraNotConnect)
}
if c.Connected == nil || *c.Connected {
t.Error("a reported failure must stay false, not become unknown")
}
}
// An unparseable timestamp is not a working camera. It should not be possible,
// which is exactly why it must not fall through to "connected".
func TestAnUnreadableTimestampIsStale(t *testing.T) {
c := Camera{Connected: ptr(true), LastSeenAt: "not a time"}
c.CameraState(time.Now())
if c.State != CameraStale || c.Connected != nil {
t.Errorf("state = %q connected = %v, want stale and unknown", c.State, c.Connected)
}
}

View File

@@ -655,6 +655,11 @@ type Camera struct {
Snapshot Image `json:"snapshot"`
SnapshotAt string `json:"snapshot_at,omitempty"`
// State is the ONE answer a screen should render, because there are four
// of them and only three were ever expressed. See CameraState.
State string `json:"state"`
StateNote string `json:"state_note,omitempty"`
// Check is the last attempt to prove this camera works. Always present so
// a client can tell "never checked" from "checked and failed" without
// guessing from an absent field.
@@ -962,3 +967,70 @@ type DeviceSession struct {
// hand.
Current bool `json:"current"`
}
// Camera states, and why a fourth one had to exist.
//
// A shop computer reports each camera's state about once a minute. When it
// cannot reach the recognition engine it reports NOTHING - correctly, because
// it has nothing to say - and the last value it sent stays in the database
// unchanged. Measured on the live estate: two cameras reading **Connected**,
// in green, thirty-four minutes after the shop computer had stopped being
// able to see either of them, while the heartbeat from the same PC said
// 0 of 0 cameras. Both surfaces were reading stored fields and disagreeing.
//
// `false` could not be the answer. It means "this camera is not connecting",
// which sends an installer to check cabling on a camera that was working
// perfectly the last time anybody could ask it. The honest statement is that
// nobody currently knows - and that is a different sentence from "nobody has
// ever told us", which is what an unreported camera needs. Two states that
// need different actions must never share a word; the same rule that keeps
// `artifact` apart from `no_faces` in the commissioning verdicts.
const (
CameraConnected = "connected" // reported recently, and working
CameraNotConnect = "not_connecting" // reported recently, and not
CameraWaiting = "waiting" // no shop computer has ever reported
CameraStale = "stale" // reported once, and not lately
)
// CameraStaleAfter is five missed reports, not one.
//
// The agent reports on a 60 s tick, so one miss is a dropped packet or a slow
// upload. Calling that stale would put a warning on a healthy estate every few
// minutes, and an indicator that cries wolf is one people learn to ignore -
// which is the same reasoning that makes a site offline after three missed
// heartbeats rather than one.
const CameraStaleAfter = 5 * time.Minute
// CameraState decides the four states, and nulls Connected when it is not
// entitled to an opinion.
//
// Connected is CLEARED rather than left alone on purpose. Leaving a stale true
// in place would keep the lie available to every client that reads the field
// directly - a mobile app, a script, an older build of our own desktop app -
// and would leave two fields on one object disagreeing, which is precisely how
// the shops screen once came out labelled Working in green above "2 of 3
// cameras not connecting".
func (c *Camera) CameraState(now time.Time) {
switch {
case c.LastSeenAt == "":
c.State = CameraWaiting
c.StateNote = "No shop computer has reported on this camera yet."
c.Connected = nil
return
}
seen, err := time.Parse(time.RFC3339, c.LastSeenAt)
if err != nil || now.Sub(seen) > CameraStaleAfter {
c.State = CameraStale
c.StateNote = "The shop computer has stopped reporting this camera. " +
"Check that the computer is on and Behavision is running on it."
c.Connected = nil
return
}
if c.Connected != nil && *c.Connected {
c.State = CameraConnected
return
}
c.State = CameraNotConnect
c.StateNote = "The shop computer cannot open this camera's stream. " +
"Check the address, the password and the cabling."
}

View File

@@ -59,6 +59,11 @@ func scanCamera(row pgx.Row) (api.Camera, error) {
// The KEY travels in ImageKey, which is json:"-", and the handler swaps it
// for a signed link. Same rule as an arrival's face.
c.Snapshot.Key = snapKey
// Here rather than in a handler, because every camera anybody reads comes
// through this function and a state computed per caller is a state one
// caller forgets - which is how a stale `connected` reached three screens
// at once.
c.CameraState(time.Now())
return c, nil
}

File diff suppressed because one or more lines are too long

View File

@@ -6,7 +6,7 @@
<meta name="color-scheme" content="dark" />
<link rel="icon" type="image/png" href="/favicon.png" />
<title>Behavision</title>
<script type="module" crossorigin src="/assets/index-Ckr5hGZd.js"></script>
<script type="module" crossorigin src="/assets/index-BAlXnnln.js"></script>
<link rel="stylesheet" crossorigin href="/assets/index-D4KGRSVS.css">
</head>
<body>

View File

@@ -79,3 +79,48 @@ export const MAKES = [
]
export const makeById = (id) => MAKES.find(m => m.id === id) || MAKES[MAKES.length - 1]
// Paste the camera's RTSP URL, rather than taking it apart by hand.
//
// The engine has always accepted a whole URL (`CameraConfig.url` wins over the
// parts) and no form has ever offered one - the same gap as the webcam option,
// and it costs more here. A URL is how people actually HAVE this information:
// it is what the camera's own app shows, what an installer writes down and
// what gets pasted into a message. Splitting it into five fields by eye is
// where a password containing `@` or `/` goes wrong, and this repository
// already records a whole class of bug from unencoded `@` in RTSP credentials.
//
// Split into fields rather than stored whole, deliberately: everything else on
// the form - Test, the make picker, editing later, and the rule that a
// password is never returned to the browser - works on the parts. A URL kept
// intact would carry the password back out to every screen that reads a
// camera.
export function parseRtspUrl(raw) {
const text = String(raw || '').trim()
if (!text) return null
const hasScheme = /^[a-z][a-z0-9+.-]*:\/\//i.test(text)
// A bare `host/path` is a reasonable thing to paste, so the scheme is
// optional - but something has to mark this as a URL rather than a word.
// Without the slash test, `nonsense` parses as a perfectly good hostname
// and silently fills the Address field with it: a wrong answer that looks
// like it worked, which is worse than refusing.
if (!hasScheme && !text.includes('/')) return null
const withScheme = hasScheme ? text : 'rtsp://' + text
let u
try {
u = new URL(withScheme)
} catch {
return null
}
if (!u.hostname) return null
// WHATWG splits user info at the LAST `@`, which is what makes an unencoded
// `@` inside a password parse the way a person means it.
const out = {
host: u.hostname,
port: Number(u.port) || 554,
path: (u.pathname || '') + (u.search || '') || '/',
username: decodeURIComponent(u.username || ''),
password: decodeURIComponent(u.password || ''),
}
return out
}

View File

@@ -225,3 +225,42 @@ def test_an_inverted_per_camera_pair_is_rejected_not_stored(client):
tuning={"enroll_threshold": 0.8, "match_threshold": 0.5})
assert r.status_code == 400
assert c.get("/api/cameras").json() == []
# "This computer's own camera" — the demo case, and a capability the engine
# has always had with nothing able to reach it.
#
# It matters most for showing the product to somebody. A laptop's own camera
# gives real recognition, of real faces, in the room, depending on no network
# at all — where pointing a demo machine at a camera in another building
# depends on two internet connections and a tunnel staying up while you talk.
def test_this_computers_own_camera_can_be_added(client):
c, eng = client
r = c.post("/api/cameras", json={"id": "laptop", "webcam": 0})
assert r.status_code == 201, r.text
assert "laptop" in eng.workers
cam = next(x for x in c.get("/api/cameras").json() if x["id"] == "laptop")
# Returned, or the form cannot tell a webcam camera from a half-filled
# RTSP one when somebody opens it to edit.
assert cam["webcam"] == 0
assert eng.camera_store.get("laptop").source() == 0
def test_a_second_camera_index_is_kept(client):
"""0 is the built-in one; a plugged-in camera is usually 1. An index
silently coerced to 0 would open the wrong camera and look like the
setting had no effect."""
c, eng = client
assert c.post("/api/cameras", json={"id": "usb", "webcam": 1}).status_code == 201
assert eng.camera_store.get("usb").source() == 1
def test_a_webcam_camera_survives_a_reload(client, tmp_path):
"""It has to be on disk, not only in the running engine: a demo that
forgets its camera when the app restarts is worse than no demo."""
c, _ = client
c.post("/api/cameras", json={"id": "laptop", "webcam": 0})
again = CameraStore(tmp_path / "cameras.json").get("laptop")
assert again.source() == 0
assert again.safe_url() == "webcam:0"

View File

@@ -1,7 +1,32 @@
"""cv2.FaceDetectorYN caches its input size and is not thread-safe, so camera
workers must not share one. Skipped when the model is absent, matching the
faiss-optional pattern in test_index.py — the suite stays runnable with no
models installed."""
models installed.
The premise is proved in a SUBPROCESS, and that is the whole lesson of this
file. The first version raced a shared detector in-process and asserted that
an exception came back, because on OpenCV 4.11 one usually did. It is a data
race in C++: what it produces is undefined, and on 4.14 what it mostly
produces is a segmentation fault. Measured, nine runs each, same machine:
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
numpy is not the variable; OpenCV is, and 4.11 crashing once says the hazard
was always there and 4.11 merely survived it. A test that takes the whole
suite down two runs in three is worse than no test: it turns "we upgraded
OpenCV" into a CI failure with no failing assertion in it, which is the
hardest kind to read.
None of this reaches the product. `Engine._build_worker` constructs a
FaceDetector per camera, which is what the rule says and what the second test
here guards.
"""
import subprocess
import sys
import textwrap
import threading
from pathlib import Path
@@ -11,10 +36,16 @@ import pytest
from behavision.config import Config
from behavision.detection import YUNET_FILENAME, FaceDetector
MODELS = Path(__file__).resolve().parent.parent / "models"
ROOT = Path(__file__).resolve().parent.parent
MODELS = ROOT / "models"
pytestmark = pytest.mark.skipif(not (MODELS / YUNET_FILENAME).exists(),
reason="YuNet model not installed")
# A shared detector fails probabilistically, so one clean attempt proves
# nothing. Several do: the premise holds if ANY attempt misbehaves, and only
# an unbroken run of clean ones is evidence it has stopped being true.
ATTEMPTS = 6
def _detector():
d = Config().detection
@@ -44,11 +75,37 @@ def _race(det_a, det_b):
return errors
_SHARED_RACE = textwrap.dedent("""
import sys
sys.path.insert(0, {root!r})
from tests.test_detector_concurrency import _detector, _race
shared = _detector()
print("raced" if _race(shared, shared) else "clean")
""")
def _shared_race_outcome():
"""Run one shared-detector race out of process.
Returns "raced" (an exception came back), "crashed" (the process died,
which is the same premise arriving by a blunter route) or "clean".
"""
proc = subprocess.run([sys.executable, "-c", _SHARED_RACE.format(root=str(ROOT))],
capture_output=True, text=True, timeout=120)
if proc.returncode != 0:
return "crashed"
return proc.stdout.strip().splitlines()[-1] if proc.stdout.strip() else "clean"
def test_a_shared_detector_really_does_race():
"""Guards the premise: if this ever stops failing, the test below is
proving nothing and the per-camera split can be revisited."""
shared = _detector()
assert _race(shared, shared), "expected a shared detector to race"
seen = [_shared_race_outcome() for _ in range(ATTEMPTS)]
assert any(o != "clean" for o in seen), (
f"a shared detector survived {ATTEMPTS} races ({seen}) - if that is "
"reproducible, cv2.FaceDetectorYN may have become thread-safe and the "
"per-camera rule in CLAUDE.md can be revisited"
)
def test_per_camera_detectors_do_not_race():

View File

@@ -0,0 +1,164 @@
"""Downloading the models is the last step of every fresh install, and it ran
into the one macOS trap nothing else here does.
A python.org macOS build ships its own OpenSSL with NO trust store, and
populates one only when somebody double-clicks `Install Certificates.command`
in the Python folder. Measured on a colleague's Mac: the engine installed
perfectly - numpy, onnxruntime, faiss, all of it - and then could not fetch a
230 KB model file, ending setup in forty lines of traceback about `_ssl.c`.
"""
import io
import logging
import os
import ssl
import urllib.request
import pytest
from behavision import model_assets
class _Resp(io.BytesIO):
"""Enough of an http response for _fetch: read() and .headers."""
def __init__(self, payload: bytes):
super().__init__(payload)
self.headers = {"Content-Length": str(len(payload))}
def __enter__(self):
return self
def __exit__(self, *exc):
self.close()
return False
def _verify_error():
"""What urlopen ACTUALLY raises, which is not what it looks like.
urllib catches ssl.SSLCertVerificationError and re-raises
urllib.error.URLError(err), carrying the original on `.reason`. The first
version of this test raised the bare SSL error - a shape real urllib never
produces - so it passed against a fallback that could never fire, and the
fix shipped and failed on the machine it was written for with the exact
traceback it was meant to prevent.
"""
return urllib.error.URLError(ssl.SSLCertVerificationError(
"[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: "
"unable to get local issuer certificate"))
def test_a_machine_with_no_trust_store_falls_back_to_the_bundled_one(monkeypatch):
calls = []
def fake(url, timeout=None, context=None):
calls.append(context)
if context is None:
raise _verify_error()
return _Resp(b"ok")
monkeypatch.setattr(urllib.request, "urlopen", fake)
with model_assets._urlopen("https://example.invalid/m.onnx") as resp:
assert resp.read() == b"ok"
assert len(calls) == 2, f"expected a retry, got {calls}"
# Default FIRST, and that order is the point. On Windows and on a system
# Python the default context reads the machine's own certificate store,
# which is what makes a corporate proxy with its own root CA work.
# Replacing it unconditionally would break every site that has one in
# order to fix a different platform.
assert calls[0] is None
assert isinstance(calls[1], ssl.SSLContext)
def test_a_working_trust_store_is_used_as_is(monkeypatch):
calls = []
def fake(url, timeout=None, context=None):
calls.append(context)
return _Resp(b"ok")
monkeypatch.setattr(urllib.request, "urlopen", fake)
model_assets._urlopen("https://example.invalid/m.onnx").close()
assert calls == [None], "the machine's own certificate store was bypassed"
def test_a_real_network_failure_is_not_disguised_as_a_certificate_problem(monkeypatch):
"""A URLError is not automatically a certificate problem.
"No route to host" and "connection refused" arrive as URLError too, and
retrying those with a different CA list changes nothing except how long
the operator waits for the real message.
"""
calls = []
def fake(url, timeout=None, context=None):
calls.append(context)
raise urllib.error.URLError("no route to host")
monkeypatch.setattr(urllib.request, "urlopen", fake)
with pytest.raises(urllib.error.URLError):
model_assets._urlopen("https://example.invalid/m.onnx")
assert calls == [None], f"a dead network was retried as a CA problem: {calls}"
def test_progress_lines_survive_the_rewrite(tmp_path, monkeypatch, caplog):
"""`download: <label> <n>%` is a CONTRACT, not logging.
The supervisor parses it (progressRe) to put first-run progress in the
tray and the window, because the API is not up yet and a shop PC showing
a stopped engine for five minutes after install looks broken. Switching
off urlretrieve - needed because it offers no way to pass an SSL context -
is exactly the kind of change that drops it silently.
"""
payload = b"x" * (64 * 1024 * 8)
monkeypatch.setattr(urllib.request, "urlopen",
lambda url, timeout=None, context=None: _Resp(payload))
dest = tmp_path / "model.onnx"
with caplog.at_level(logging.INFO, logger="behavision.model_assets"):
model_assets._fetch("https://example.invalid/m.onnx", dest, "face detector")
assert dest.read_bytes() == payload
lines = [r.getMessage() for r in caplog.records]
pct = [ln for ln in lines if ln.startswith("download: face detector ")]
assert pct, f"no progress lines at all: {lines}"
assert "download: face detector 100%" in pct, pct
@pytest.mark.skipif(not os.environ.get("BEHAVISION_NETWORK_TESTS"),
reason="set BEHAVISION_NETWORK_TESTS=1 to reach github.com")
def test_against_a_real_machine_with_no_trust_store(tmp_path, monkeypatch):
"""The whole thing, over the real network, with the real failure.
Every stub above is a statement about what I believe urllib does, and the
first version of this file proved how much that is worth. Python's default
context honours SSL_CERT_FILE, so pointing it at an EMPTY file reproduces
a python.org macOS build exactly: a default context that trusts nobody.
certifi is loaded by path and is unaffected by the variable, so if the
fallback works here it works there.
"""
empty = tmp_path / "no-cas.pem"
empty.write_text("")
monkeypatch.setenv("SSL_CERT_FILE", str(empty))
monkeypatch.setenv("SSL_CERT_DIR", str(tmp_path / "nothing"))
# Not every interpreter can be put into that state, and pretending
# otherwise would turn this into a test that passes by not running. A
# Python linked against LibreSSL - which is what macOS Command Line Tools
# ships - reads the system keychain and ignores SSL_CERT_FILE entirely:
# measured, 128 CAs loaded with the variable pointing at an empty file. A
# python.org build links OpenSSL and honours it: 0 CAs, which is the
# condition being reproduced.
if ssl.create_default_context().get_ca_certs():
pytest.skip("this interpreter ignores SSL_CERT_FILE (%s), so the "
"no-trust-store condition cannot be reproduced here"
% ssl.OPENSSL_VERSION)
# The premise: the default context really is broken in this process.
with pytest.raises(urllib.error.URLError) as caught:
urllib.request.urlopen(model_assets.YUNET_URL, timeout=30)
assert model_assets._is_cert_failure(caught.value), caught.value
dest = tmp_path / "yunet.onnx"
model_assets._fetch(model_assets.YUNET_URL, dest, "face detector")
assert dest.stat().st_size > 200_000

91
tests/test_rtsp_paste.py Normal file
View File

@@ -0,0 +1,91 @@
"""`shared/cameraMakes.js` is imported by BOTH camera forms and has no test
runner of its own. This is the same safety net `test_dashboard.py` provides:
node is driven from pytest and the check is skipped when node is absent, so
the suite stays dependency-light.
What it guards is the field people actually have. The engine has always
accepted a whole RTSP URL (`CameraConfig.url` wins over the parts) and no form
ever offered one, so an operator holding the address their camera's own app
shows had to take it apart into five fields by eye — which is exactly where a
password containing `@` goes wrong, a class of bug this repository has already
been bitten by once.
"""
import json
import shutil
import subprocess
from pathlib import Path
import pytest
ROOT = Path(__file__).resolve().parent.parent
SHARED = ROOT / "shared" / "cameraMakes.js"
pytestmark = pytest.mark.skipif(shutil.which("node") is None,
reason="node not installed")
def parse(text):
out = subprocess.run(
["node", "--input-type=module", "-e",
f"import {{ parseRtspUrl }} from {json.dumps(str(SHARED))};"
f"console.log(JSON.stringify(parseRtspUrl({json.dumps(text)})))"],
capture_output=True, text=True, timeout=60)
assert out.returncode == 0, out.stderr
return json.loads(out.stdout.strip())
def test_a_plain_url_becomes_the_five_fields():
assert parse("rtsp://admin:Pass123@192.168.1.121:554/ch0_0.264") == {
"host": "192.168.1.121", "port": 554, "path": "/ch0_0.264",
"username": "admin", "password": "Pass123"}
def test_an_at_sign_in_the_password_survives():
"""The one that matters. A URL is split at the LAST `@`, which is what
makes an unencoded `@` inside a password parse the way a person means it -
and an operator splitting this by eye would put `p` in the password box
and `ssw0rd@192.168.1.121` in the address box."""
got = parse("rtsp://admin:p@ssw0rd@192.168.1.121:554/Streaming/Channels/101")
assert got["password"] == "p@ssw0rd"
assert got["host"] == "192.168.1.121"
def test_a_percent_encoded_password_is_decoded():
"""Stored decoded, because CameraConfig.source() percent-encodes when it
rebuilds the URL. Keeping it encoded would double-encode it and the camera
would refuse a password that is correct."""
assert parse("rtsp://admin:p%40ss@10.0.0.5/live")["password"] == "p@ss"
def test_the_default_port_is_filled_in():
assert parse("rtsp://192.168.1.122/ch0_1.264")["port"] == 554
def test_a_query_string_stays_with_the_path():
"""Some cameras carry the channel in a query. Dropping it opens the wrong
channel, which looks like a camera pointed somewhere unexpected."""
assert parse("rtsp://cam.local:8554/live?channel=1")["path"] == "/live?channel=1"
def test_a_scheme_is_optional_but_structure_is_not():
assert parse("192.168.1.122/ch0_1.264")["host"] == "192.168.1.122"
# A bare word parses as a perfectly good hostname, so without this it
# would silently fill the Address field with it - a wrong answer that
# looks like it worked, which is worse than refusing.
assert parse("nonsense") is None
assert parse("192.168.1.121") is None, "a bare address is not a URL; the Address field takes it"
def test_nothing_is_nothing():
assert parse("") is None
assert parse(" ") is None
def test_both_forms_import_it():
"""Two copies of this would be worse than not offering it, because an
operator trusts a filled-in field. Same rule as the make picker."""
for form in (ROOT / "desktop/frontend/src/views/Cameras.jsx",
ROOT / "web/src/views/CameraSetup.jsx"):
src = form.read_text(encoding="utf-8")
assert "parseRtspUrl" in src, f"{form.name} does not offer the paste field"
assert "cameraMakes.js" in src, f"{form.name} defines its own parser"

View File

@@ -0,0 +1,81 @@
"""Why a camera "will not connect", when the real answer is that the computer
asking is in the wrong building.
Asked directly by the owner, about his own cameras, from his phone's
connection: *"when i connect from my mobile internet the cameras wont connect,
why is that"*. The answer was `cannot reach 192.168.1.121:554 - Operation
timed out`, which reads as a broken camera and sends somebody to re-type an
address and a password that were always correct.
"""
import pytest
from behavision import capture
@pytest.fixture
def on(monkeypatch):
def _set(ip):
monkeypatch.setattr(capture, "_local_ipv4", lambda: ip)
return _set
def test_a_different_network_says_so_and_says_what_to_do(on):
on("10.11.12.13") # a phone's tethered network
hint = capture._wrong_network_hint("192.168.1.121")
assert "not the camera's network" in hint
# The action, not just the diagnosis: no setting on this screen fixes it.
assert "computer in the shop" in hint
def test_the_same_network_sends_you_to_the_camera_instead(on):
on("192.168.1.120")
hint = capture._wrong_network_hint("192.168.1.121")
assert "on that network" in hint
assert "powered on" in hint
# Two states that need opposite actions must not share a sentence.
assert "not the camera's network" not in hint
@pytest.mark.parametrize("host", [
"8.8.8.8", # plainly routable
"203.0.113.9", # TEST-NET-3: `is_private` calls this private, and it is
# not a LAN address - which is why the check spells out
# the RFC1918 blocks instead of asking is_private
"100.64.0.5", # carrier-grade NAT, what a mobile network hands out
])
def test_an_address_that_is_not_a_lan_address_gets_no_hint(on, host):
"""A routable address unreachable from here is an ordinary network fault,
and inventing a story about private networks would be wrong."""
on("192.168.1.120")
assert capture._wrong_network_hint(host) == ""
def test_a_name_gets_no_hint(on):
"""Nothing can be concluded about `camera.local` from the string, and a
guess here is a confident wrong answer in the place people look first."""
on("192.168.1.120")
assert capture._wrong_network_hint("camera.local") == ""
def test_loopback_gets_no_hint(on):
on("192.168.1.120")
assert capture._wrong_network_hint("127.0.0.1") == ""
def test_with_no_network_at_all_it_still_names_the_cause(on):
"""A machine with no route cannot say which network it is on, and must not
pretend: the private-address fact is still true and still the reason."""
on("")
hint = capture._wrong_network_hint("192.168.1.121")
assert "private address" in hint
assert "this computer is on" not in hint
def test_the_hint_reaches_the_message_a_person_reads():
"""The whole point is the sentence on the screen, not a helper nobody
calls. Port 1 on a private address refuses or times out immediately."""
ok, msg = capture._tcp_reachable("rtsp://192.168.1.121:1/ch0", 1.0)
assert not ok
assert "192.168.1.121" in msg
# Whichever branch the OS takes, the explanation travels with it.
assert "network" in msg

View File

@@ -1,6 +1,6 @@
import { useEffect, useRef, useState } from 'react'
import { api } from '../api.js'
import { MAKES, makeById } from '../../../shared/cameraMakes.js'
import { MAKES, makeById, parseRtspUrl } from '../../../shared/cameraMakes.js'
// Setting up a camera, for somebody who has never done it.
//
@@ -30,6 +30,25 @@ export default function CameraSetup({ sites, existing, onClose, onSaved }) {
const [error, setError] = useState('')
const set = (k) => (e) => setForm(f => ({ ...f, [k]: e.target.value }))
// Paste the whole RTSP address. It is how people actually hold this
// information - it is what the camera's own app shows and what an installer
// writes down - and splitting it into five fields by eye is where a
// password containing `@` or `/` goes wrong.
const [pasted, setPasted] = useState('')
const [pasteError, setPasteError] = useState('')
const applyUrl = (text) => {
setPasted(text)
if (!text.trim()) { setPasteError(''); return }
const got = parseRtspUrl(text)
if (!got) { setPasteError('That does not look like an RTSP address.'); return }
setPasteError('')
// Only what the URL actually carried: one with no credentials must not
// wipe a password already typed.
setForm(f => ({ ...f, make: 'manual', host: got.host, port: got.port, path: got.path,
...(got.username ? { username: got.username } : {}),
...(got.password ? { password: got.password } : {}) }))
}
const chooseMake = (e) => {
const m = makeById(e.target.value)
// Only overwrite the path when the preset has one, so choosing "I know the
@@ -115,6 +134,14 @@ export default function CameraSetup({ sites, existing, onClose, onSaved }) {
{step === 1 && (
<div className="drawer-body">
<label>Paste the camera’s RTSP address, if you have one
<input className="mono" value={pasted} onChange={(e) => applyUrl(e.target.value)}
placeholder="rtsp://admin:password@192.168.0.138:554/ch0_0.264"
autoComplete="off" name="rtsp-url" spellCheck="false" />
<span className="hint">{pasteError
? pasteError
: 'Optional. Paste it and the fields below fill in; otherwise fill them in yourself.'}</span>
</label>
<label>Camera’s address on the shop’s network
<input value={form.host} onChange={set('host')}
placeholder="192.168.0.138" autoFocus />

View File

@@ -47,9 +47,13 @@ export default function Cameras({ user }) {
const canEdit = ['admin', 'owner', 'manager'].includes(user.role)
const list = cams || []
const up = list.filter(c => c.connected).length
const down = list.filter(c => c.connected === false).length
const waiting = list.filter(c => c.connected == null).length
// Counted off `state`, the one field the server computes, never off
// `connected`. Two places deciding the same fact is how a shop came out
// labelled Working, in green, above "2 of 3 cameras not connecting".
const up = list.filter(c => c.state === 'connected').length
const down = list.filter(c => c.state === 'not_connecting').length
const waiting = list.filter(c => c.state === 'waiting').length
const stale = list.filter(c => c.state === 'stale').length
return (
<>
@@ -58,6 +62,7 @@ export default function Cameras({ user }) {
<p className="sub">
{list.length} {list.length === 1 ? 'camera' : 'cameras'}
{up > 0 && <> · <b className="ok">{up} connected</b></>}
{stale > 0 && <> · <b className="warn">{stale} not reporting</b></>}
{down > 0 && <> · <b className="bad">{down} down</b></>}
{waiting > 0 && <> · {waiting} waiting for the shop PC</>}
</p>
@@ -108,10 +113,13 @@ function CameraCard({ cam, canEdit, onEdit, onWatch }) {
// Three states, not two. A camera nobody has tried yet is not a camera that
// is down, and telling an operator to check the cabling on a camera the shop
// PC has not even seen sends them to the wrong building.
const state = cam.connected == null ? 'idle'
: cam.connected ? 'ok' : 'bad'
const words = cam.connected == null ? 'Waiting for the shop PC'
: cam.connected ? 'Connected' : 'Not connecting'
// Four states, and the fourth is the one that was missing: a camera whose
// shop PC has stopped reporting it. The agent correctly says nothing when it
// cannot reach the engine, so the last value it sent used to sit in the
// database reading Connected - measured at 34 minutes on the live estate.
const state = { connected: 'ok', not_connecting: 'bad', stale: 'warn' }[cam.state] || 'idle'
const words = { connected: 'Connected', not_connecting: 'Not connecting',
stale: 'Not reporting' }[cam.state] || 'Waiting for the shop PC'
const verified = verification(cam)
return (