4 Commits

Author SHA1 Message Date
b296e8a74a A demo does not need a tunnel; it needs the laptop's own camera
Asked after the mobile-internet question: could a VPN let the office cameras
be shown in the demo. Three jobs get confused there and only one needs one.

Seeing the estate from anywhere already works and needs nothing - viewer mode
plus LiveHub is exactly that, outbound, no installation and no credential.

Demonstrating recognition is better done on the demo machine's own camera.
`webcam: 0` picks a capture index instead of building an RTSP URL and the
engine has supported it since the first version: CameraStore round-trips it,
source() returns the index, safe_url() reports webcam:0, and
POST /api/cameras {"id":"laptop","webcam":0} has always worked. No screen
offered it - the same gap this repo already records for the customer record
and per-camera tuning. It is now an option in the make picker, and it is the
strongest demo available: real faces, in the room, depending on no network.
A demo pointed at a camera in another building depends on two internet
connections and a tunnel staying up while somebody is talking.

The address and the index are alternatives, not extras: source() takes the
webcam first, so a half-typed host left behind would make the saved camera
describe two things and use one. The scan, the path and the camera password
are hidden for a local camera because none of them mean anything.

A tunnel is still right for one case - running the engine on a remote machine
against the office's own cameras - and still wrong for the product: it is
per-site infrastructure on every shop PC, and it gives head office
network-level access into a customer's LAN, where today we can read a
camera's picture and nothing else.

Also: installing httpx took the suite from 239 passed to 271. The HTTP tests
importorskip it so a bare checkout runs, which means the number at the bottom
of a run is not the number of tests that exist. Added to the dev extra.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 18:18:57 +05:30
c9be9b5807 "The cameras won't connect from my mobile internet" was answered by a timeout
Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics and the product was not giving it.

A camera lives on the shop's LAN behind a router. 192.168.1.121 means
"something on the network I am attached to" and nothing more - from mobile
data, a hotel or head office it resolves to nobody, or to a completely
different device holding that number. There is no route in from the internet
and there must not be: an RTSP camera reachable from outside is how a shop's
cameras end up being watched by strangers.

That is why the product is split the way it is - the shop PC is the only
machine on the camera's LAN, and every other surface reaches it outbound,
which is what makes Watch live work from anywhere while nothing connects in.

What was wrong is the message. "cannot reach 192.168.1.121:554 - Operation
timed out" reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. _wrong_network_hint names the cause
and separates two states that need opposite actions:

  on that network   -> check the camera is powered on and the address is right
  somewhere else    -> the COMPUTER is in the wrong place; no setting fixes it

- The local address comes from a connected UDP socket that sends nothing. It
  only fixes a route so the kernel will name the source address.
- The LAN ranges are spelled out, not is_private. That property also covers
  carrier-grade NAT and the documentation networks, and telling somebody who
  typed 203.0.113.9 that it is "on the shop's own network" is a confident
  wrong answer in the place people look first. Found by a test using that
  address as its example of a PUBLIC one.
- A DNS name gets no hint: nothing can be concluded about camera.local from
  the string, and guessing is the failure mode this message exists to fix.
- With no network at all it still names the cause and drops the comparison,
  rather than claiming to know which network this machine is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 18:00:44 +05:30
0558344dc2 urllib does not let SSLCertVerificationError out, and a fake said it did
The certificate fix shipped and failed on the machine it was written for, with
the exact traceback it was meant to prevent. The retry was written

    except ssl.SSLCertVerificationError:

and urllib never raises that from urlopen. It catches it and re-raises
urllib.error.URLError(err), carrying the original on .reason. So the except
matched nothing, ever, and the fallback could not fire.

The unit test passed throughout, because the stub it used raised the bare SSL
error - a shape real urllib never produces. That is the lesson: a fake that
agrees with the author is worse than no test, because it converts an untested
path into a tested-looking one. This file already says that about
UPDATE ... RETURNING and about the in-memory API fake, and it got written
again anyway.

_is_cert_failure checks the exception and its .reason, and the tests now raise
URLError(SSLCertVerificationError(...)) - what the traceback actually shows. A
plain URLError is re-raised untouched, and a test asserts no second attempt is
made for one.

Beside the stubs there is now a real reproduction, opt-in behind
BEHAVISION_NETWORK_TESTS=1. Python's default context honours SSL_CERT_FILE, so
an empty file gives a context that trusts nobody - the python.org condition
exactly - while certifi is loaded by path and is unaffected. It skips rather
than passes where it cannot reproduce that, and the difference is measured:

    macOS Command Line Tools  LibreSSL 2.8.3   128 CAs with an empty CA file
    python.org / pyenv build  OpenSSL 3.5.8      0 CAs -> reproduces it

Checked for teeth by putting the shipped except back: both the corrected stub
test and the live one fail, and pass again when it is restored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:36:53 +05:30
8f07dee048 The engine stopped updating, and said "installed" every time
The SSL fix shipped and did not reach the machine it was written for. That
log said so, one line above the tick:

  behavision is already installed with the same version as the provided
  wheel. Use --force-reinstall to force an installation of the wheel.
   [ok] Engine and dependencies installed

and the traceback below it still pointed at model_assets.py line 59,
urllib.request.urlretrieve - code the fix had deleted.

Two frozen literals caused it: version = "1.1.0" in pyproject.toml and
__version__ = "1.0.0" in behavision/__init__.py. They disagreed with each
other and neither tracked a release, so every release built
behavision-1.1.0-py3-none-any.whl and pip install --upgrade on a machine that
already had 1.1.0 is a no-op. The comment beside that call claimed the
opposite.

The shape of the damage is what makes it bad. The Go binaries - app, agent,
setup tool - are rebuilt every release and updated normally. So a shop PC ran
a current app supervising an engine several releases old, and nothing said
which: /api/health reported the model, the paths, the cameras and the gallery,
and no version at all.

- One version, in the package, read by pyproject through
  [tool.setuptools.dynamic]. In a checkout it reads 0.0.0+dev: a
  plausible-looking number on a developer's /api/health is worse than none.
- pip install --force-reinstall --no-deps <wheel>, after the ordinary
  --upgrade. --upgrade settles the dependencies; the second call guarantees our
  own code is the code in the folder. --no-deps keeps it cheap - forcing the
  dependencies too would re-download ~300 MB every run. A rebuild at an
  unchanged version is the ordinary case while developing, so this must not
  rely on the version moving.
- release.sh stamps the tag: v0.5.6-demo -> 0.5.6+demo, valid PEP 440. Into a
  copy of the line, reverted in a trap, so the tree is never left dirty.

/api/health reports version now. Without it there is no way to tell a shop PC
three releases behind from a current one, which is how this survived several
releases.

Reproduced end to end before fixing, against real wheels on Python 3.12: two
builds of the same version, --upgrade leaves the old code in place and prints
the same sentence the colleague's Mac printed, --force-reinstall --no-deps
replaces it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:30:24 +05:30
15 changed files with 668 additions and 80 deletions

203
CLAUDE.md
View File

@@ -3599,3 +3599,206 @@ is not up yet and a shop PC showing a stopped engine for five minutes after
install looks broken. `tests/test_model_download.py` asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.
## The engine stopped updating, and said "installed" every time
The SSL fix above was released and **did not reach the machine it was written
for**. Its log said so, one line above the tick:
```
behavision is already installed with the same version as the provided wheel.
Use --force-reinstall to force an installation of the wheel.
[ok] Engine and dependencies installed
```
and the traceback that followed still pointed at `model_assets.py", line 59,
in _fetch / urllib.request.urlretrieve` — code the fix had deleted.
Two frozen literals caused it: `version = "1.1.0"` in `pyproject.toml` and
`__version__ = "1.0.0"` in `behavision/__init__.py`. They disagreed with each
other and neither tracked a release, so **every release built
`behavision-1.1.0-py3-none-any.whl`** and `pip install --upgrade` on a machine
that already had 1.1.0 is a no-op. The comment beside that call even claimed
the opposite: *"`--upgrade` so re-running after a new release replaces the
engine rather than leaving the old one in place and reporting success."*
The shape of the damage is what makes it bad. The Go binaries — the app, the
agent, the setup tool — are rebuilt every release and updated normally. So a
shop PC ran a current app supervising an engine several releases old, and
nothing anywhere said which: `/api/health` reported the model, the paths, the
cameras and the gallery, and **no version at all**.
Three changes, and the second is the one that does not depend on the first
being remembered:
- **One version, in the package**, read by `pyproject.toml` through
`[tool.setuptools.dynamic]`. Two literals in two files is how they came to
disagree. In a checkout it reads `0.0.0+dev`, because a plausible-looking
number on a developer's `/api/health` would be worse than none.
- **`pip install --force-reinstall --no-deps <wheel>`, after the ordinary
`--upgrade`.** `--upgrade` settles the dependencies; the second call
guarantees our own code is the code in the folder. `--no-deps` is what keeps
it cheap — forcing the dependencies too would re-download ~300 MB of numpy,
OpenCV and onnxruntime on every run. A rebuild at an unchanged version is the
ordinary case while developing, so this must not rely on the version moving.
- **`release.sh` stamps the tag**: `v0.5.6-demo` → `0.5.6+demo`, valid PEP 440
(a local segment takes alphanumerics and dots, never hyphens). Into a copy of
the line, reverted in a `trap`, so the tree is never left dirty.
`/api/health` reports `version` now. Without it there is no way to tell a shop
PC three releases behind from a current one, which is precisely how this
survived several releases.
Reproduced end to end before fixing, against real wheels on Python 3.12: two
builds of the same version, `--upgrade` leaves the old code in place and prints
the same sentence the colleague's Mac printed, `--force-reinstall --no-deps`
replaces it.
### `urllib` does not let `SSLCertVerificationError` out, and a fake said it did
The certificate fix above shipped and **failed on the machine it was written
for, with the exact traceback it was meant to prevent**. The retry was written
```python
except ssl.SSLCertVerificationError:
```
and urllib never raises that from `urlopen`. It catches it and re-raises
`urllib.error.URLError(err)`, carrying the original on `.reason`. So the
`except` matched nothing, ever, and the fallback could not fire.
The unit test passed throughout, because the stub it used raised the bare SSL
error — **a shape real urllib never produces**. That is the whole lesson: a
fake that agrees with the author is worse than no test, because it converts an
untested path into a tested-looking one. The same sentence is already in this
file about `UPDATE ... RETURNING` and about the in-memory API fake, and it was
written again here anyway.
`_is_cert_failure` checks the exception **and** its `.reason`, and the tests
now raise `URLError(SSLCertVerificationError(...))`, which is what his
traceback shows. A plain `URLError` — "no route to host" — is re-raised
untouched: retrying that with a different CA list changes nothing except how
long the operator waits for the real message, and a test asserts the second
attempt never happens.
Beside the stubs there is now a **real** reproduction,
`test_against_a_real_machine_with_no_trust_store`, opt-in behind
`BEHAVISION_NETWORK_TESTS=1` because it reaches github.com. Python's default
context honours `SSL_CERT_FILE`, so pointing it at an empty file gives a
default context that trusts nobody — the python.org condition exactly — while
certifi is loaded by path and is unaffected.
It skips rather than passes where it cannot reproduce that, and the difference
is measured rather than assumed:
```
macOS Command Line Tools LibreSSL 2.8.3 128 CAs with SSL_CERT_FILE empty
(reads the system keychain)
python.org / pyenv build OpenSSL 3.5.8 0 CAs -> reproduces it
```
Checked for teeth by putting the shipped `except` back: both the corrected stub
test and the live one fail, and pass again when it is restored.
## "The cameras won't connect from my mobile internet" — and why that is not a bug
Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics, and the product was not giving it.
A camera lives on the shop's LAN behind a router. `192.168.1.121` means
*"something on the network I am attached to"* and nothing more — from mobile
data, a hotel, or head office it resolves to nobody, or to a completely
different device that happens to hold that number. There is no route in from
the internet and **there must not be**: an RTSP camera reachable from outside
is how a shop's cameras end up being watched by strangers.
This is the reason the product is split the way it is, and it is worth stating
plainly next to the split itself: the shop PC is the only machine on the
camera's LAN, so it does the connecting, and every other surface reaches it
**outbound** — the agent's pull, the arrivals feed, and `LiveHub`'s frame relay,
which is what makes **Watch live** work from anywhere while nothing connects in.
What was wrong is the message. `cannot reach 192.168.1.121:554 - Operation
timed out` reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. `_wrong_network_hint` now names the
cause, and it distinguishes two states that need opposite actions — the same
rule as `artifact` against `no_faces`, and `stale` against `not_connecting`:
| | |
|---|---|
| this computer **is** on that network | check the camera is powered on and that the address is right |
| this computer is **somewhere else** | the computer is in the wrong place; no setting here fixes it, recognition has to run on a machine in the shop |
- **The local address comes from a `connect`ed UDP socket** that sends nothing.
It only fixes a route so the kernel will name the source address — no packet
leaves, and it needs no dependency in an engine that already ships 200 MB of
models.
- **The LAN ranges are spelled out, not `is_private`.** That property is
broader than "an address on somebody's LAN": it also covers carrier-grade NAT
and the documentation networks (192.0.2, 198.51.100, 203.0.113), and telling
somebody who typed one of those that it is "on the shop's own network" is a
confident wrong answer in the place people look first. Found by a test using
`203.0.113.9` as its example of a *public* address, which `is_private` calls
private.
- **A DNS name gets no hint at all.** Nothing can be concluded about
`camera.local` from the string, and guessing is the failure mode this whole
message exists to fix.
- **With no network at all it still names the cause** and drops the comparison,
rather than claiming to know which network this machine is on.
## A tunnel for the demo: when it is the right answer, and when it is not
Asked after the mobile-internet question: *"what if we created a secure tunnel
or vpn, then the cameras in our office can be shown in the demo version too?"*
Three different jobs get confused here, and only one of them needs a tunnel.
**Seeing the estate from anywhere already works and needs nothing.** Viewer
mode plus `LiveHub` is exactly this: the shop PC pushes frames outbound and any
signed-in app sees them, from mobile data, a hotel or a customer's office. A
VPN would add an installation, a credential and a moving part to something that
already works with none of them.
**Demonstrating recognition is better done on the demo machine's own camera.**
`webcam: 0` picks a capture index instead of building an RTSP URL, and the
engine has supported it since the first version — `CameraStore` round-trips it,
`source()` returns the index, `safe_url()` reports `webcam:0`, and
`POST /api/cameras {"id":"laptop","webcam":0}` has always worked. **No screen
offered it**, which is the same gap this file already records for the customer
record and for per-camera tuning: the API could, the UI could not reach it.
It is now an option in the make picker, and it is the strongest demo available:
real faces, in the room, instant, depending on no network at all. A demo
pointed at a camera in another building depends on two internet connections and
a tunnel staying up while somebody is talking.
**The one case a tunnel genuinely answers** is running the engine on a remote
machine against the office's own cameras — LAN access to `192.168.1.121` from
somewhere that is not that LAN. A mesh VPN (Tailscale and similar) does that
honestly: a subnet router at the office, the client on the demo machine, and
the address works unchanged. Free at this scale, no port forwarding, no
exposed camera. Worth using for our *own* office when that is really the goal.
**It is still the wrong answer for the product**, and the reasons are not about
difficulty:
- It is per-site infrastructure — an account, a node, a key — on every shop PC
we ship, to replace something that already works over ordinary HTTPS.
- It gives head office **network-level access into a customer's LAN**. The
current design can read a camera's picture; a VPN can reach the customer's
till, their router and everything else on that network. That is a far larger
thing to be trusted with, and a far larger thing to have breached.
- The failure modes are worse and less legible: a tunnel that is down looks
like a camera that is down.
So: tunnel for our own office if we want the engine running remotely; nothing
at all for showing customers their shops; the laptop's own camera for showing
anybody what the product does.
### And the suite was quietly 32 tests smaller than it looked
Installing `httpx` to check the webcam path took the engine suite from **239
passed to 271**. `tests/test_api_cameras.py` and its siblings begin with
`pytest.importorskip("httpx")` so a bare checkout still runs — deliberate, and
it means the number at the bottom of a run is not the number of tests that
exist. `pip install -e .[dev]` is the opt-in.

View File

@@ -539,7 +539,30 @@ func pipInstall(vpy, src string) error {
// 'behavision.egg-info': Read-only file system". Falling back to source
// copies it somewhere writable first, for the same reason.
if wheels, _ := filepath.Glob(filepath.Join(src, "behavision-*.whl")); len(wheels) > 0 {
return stream(exec.Command(vpy, "-m", "pip", "install", "--upgrade", wheels[0]),
// TWO calls, and the second is the one that matters.
//
// `--upgrade` alone is not an upgrade when the version has not moved:
// pip skips the wheel and says so, one line above this program
// printing "[ok] Engine and dependencies installed". The wheel version
// was a frozen literal for several releases, so every engine fix in
// them silently failed to reach any machine that had run setup once -
// while the Go binaries beside it, rebuilt every release, updated
// normally. Half the product current, half of it months old, and
// nothing saying which.
//
// release.sh stamps the tag into the version now, so the versions do
// differ. This does not rely on that: a rebuild at the same version is
// the ordinary case while developing, and "installed" has to mean the
// code in this folder either way.
if err := stream(exec.Command(vpy, "-m", "pip", "install", "--upgrade", wheels[0]),
"installing the engine"); err != nil {
return err
}
// --no-deps so this is our own package only: the call above has
// already settled the dependencies, and forcing those too would
// re-download ~300 MB of numpy, OpenCV and onnxruntime every run.
return stream(exec.Command(vpy, "-m", "pip", "install",
"--force-reinstall", "--no-deps", wheels[0]),
"installing the engine")
}
tmp, err := os.MkdirTemp("", "behavision-src-")

View File

@@ -4,4 +4,15 @@ Pipeline: capture -> detect (YuNet) -> track (IoU) -> align + encode
(ArcFace ONNX) -> match / auto-enroll (FAISS + SQLite) -> events + API.
"""
__version__ = "1.0.0"
# Stamped by release.sh from the git tag, into a copy of this line, so a
# release always produces a wheel nobody has installed before. It was a frozen
# "1.0.0" here and a frozen "1.1.0" in pyproject.toml - two literals that
# disagreed with each other and tracked nothing - which is how every engine fix
# for several releases silently failed to reach a machine that had run setup
# once. pip skips a wheel whose version is already installed and says so, one
# line above setup printing "[ok] Engine and dependencies installed".
#
# In a checkout it stays obviously a checkout: "0.0.0+dev" on /api/health is
# the truth about a developer's machine, and a plausible-looking number there
# would be worse than none.
__version__ = "0.0.0+dev"

View File

@@ -17,6 +17,7 @@ from fastapi.responses import HTMLResponse, Response, StreamingResponse
from fastapi.security import HTTPBasic, HTTPBasicCredentials
from pydantic import BaseModel, ValidationError
from . import __version__
from .config import ApiSection, CameraConfig, CameraTuning
from .commission import CommissionRun
from .events import Event
@@ -142,7 +143,7 @@ def _reencode(jpeg: bytes, width: int, quality: int) -> "bytes | None":
def create_app(engine: Engine) -> FastAPI:
app = FastAPI(title="Behavision", version="1.0.0",
app = FastAPI(title="Behavision", version=__version__,
dependencies=_auth_dependencies(engine.cfg.api))
def worker_or_404(camera_id: str):
@@ -172,6 +173,10 @@ def create_app(engine: Engine) -> FastAPI:
# are up, and the shop recognises nobody it already knows.
stranded = engine.gallery.health["stranded"]
return {"status": "ok" if engine.started_at else "starting",
# Which build is actually running. Without it there was no way
# to tell a shop PC three releases behind from a current one -
# which is precisely how a silent install failure survived.
"version": __version__,
"recognition_model": engine.encoder.model_name,
"gallery_unreadable_embeddings": stranded,
# "where is my database" must be answerable from the API: the

View File

@@ -68,6 +68,78 @@ os.environ.setdefault(
STALL_AFTER_S = 10.0
def _local_ipv4() -> str:
"""This machine's address on the interface holding the default route.
A UDP socket is `connect`ed and nothing is sent - it only fixes a route so
the kernel will name the source address. No packet leaves, and it needs no
dependency, which matters in an engine that already ships 200 MB of models.
"""
import socket
try:
with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as s:
s.settimeout(0.5)
s.connect(("8.8.8.8", 80))
return s.getsockname()[0]
except OSError:
return ""
def _wrong_network_hint(host: str) -> str:
"""Why a private camera address is unreachable, when that is the reason.
A camera lives on the shop's LAN behind a router, and 192.168.x.x means
"something on the network I am attached to" - nothing more. From mobile
data, a hotel, or head office it either resolves to nobody or to a
completely different device that happens to hold that number. There is no
route in from the internet and there must not be: an RTSP camera reachable
from outside is how a shop's cameras end up being watched by strangers.
Without this the answer was "cannot reach 192.168.1.121:554 - Operation
timed out", which reads as a broken camera and sends somebody to re-type
an address and a password that were always correct. Asked directly by the
owner, about his own cameras, from his phone's connection.
Two states, two different actions, so they must not share a sentence: on
the same network the camera or its address is the problem; on a different
one the COMPUTER is in the wrong place and no setting will fix it.
"""
import ipaddress
try:
addr = ipaddress.ip_address(host)
except ValueError:
return "" # a DNS name; nothing can be concluded from the string
# The RFC1918 blocks and link-local, spelled out rather than `is_private`.
# That property is broader than "an address on somebody's LAN": it also
# covers the carrier-grade NAT range and the documentation networks
# (192.0.2, 198.51.100, 203.0.113), and telling somebody who typed one of
# those that it is "on the shop's own network" would be a confident wrong
# answer in the place people look first. Found by a test using 203.0.113.9
# as an example of a PUBLIC address, which `is_private` calls private.
lan = any(addr in ipaddress.ip_network(n) for n in
("10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "169.254.0.0/16")
if addr.version == 4)
if not lan:
return ""
mine = _local_ipv4()
if not mine:
return (" - that is a private address, reachable only from inside "
"the network the camera is on")
try:
same = ipaddress.ip_network(f"{mine}/24", strict=False).supernet_of(
ipaddress.ip_network(f"{host}/24", strict=False))
except (ValueError, TypeError):
same = False
if same:
return (f" - this computer is on that network ({mine}), so check the "
f"camera is powered on and that {host} is its address")
return (f" - this computer is on {mine}, not the camera's network. A "
f"private address like {host} is only reachable from inside the "
f"shop's own network, never over the internet or mobile data, so "
f"recognition has to run on a computer in the shop")
def _tcp_reachable(source: "str | int", timeout: float
) -> "tuple[bool, str]":
"""Cheap pre-flight for an rtsp:// URL. Non-URL sources pass through."""
@@ -85,10 +157,11 @@ def _tcp_reachable(source: "str | int", timeout: float
return True, ""
except socket.timeout:
return False, (f"no response from {parsed.hostname}:{port} within "
f"{timeout:.0f}s - check the IP address and that the "
f"camera is on the same network")
f"{timeout:.0f}s{_wrong_network_hint(parsed.hostname)}")
except OSError as exc:
return False, f"cannot reach {parsed.hostname}:{port} - {exc.strerror or exc}"
return False, (f"cannot reach {parsed.hostname}:{port} - "
f"{exc.strerror or exc}"
f"{_wrong_network_hint(parsed.hostname)}")
def _fourcc(cap) -> str:

View File

@@ -5,6 +5,7 @@ from __future__ import annotations
import logging
import shutil
import ssl
import urllib.error
import urllib.request
from pathlib import Path
@@ -65,6 +66,27 @@ def _https_context() -> "ssl.SSLContext | None":
return ssl.create_default_context(cafile=certifi.where())
def _is_cert_failure(err: BaseException) -> bool:
"""Is this a certificate-verification failure, however it is wrapped?
`urllib` does NOT let `ssl.SSLCertVerificationError` out. It catches it and
re-raises `urllib.error.URLError(err)`, carrying the original on `.reason`
- so `except ssl.SSLCertVerificationError` around `urlopen` matches
nothing, ever.
That is not a subtlety this file gets to record academically: the first
version of the fallback below was written exactly that way, shipped, and
failed on the machine it was written for with the very traceback it was
meant to prevent. The unit test passed throughout, because the fake it
used raised the bare SSL error - a shape real urllib never produces. A
stub that agrees with the author is worse than no test, and the test now
raises what urllib raises.
"""
reason = getattr(err, "reason", None)
return isinstance(err, ssl.SSLCertVerificationError) or \
isinstance(reason, ssl.SSLCertVerificationError)
def _urlopen(url: str, timeout: float = 60.0):
"""Open a URL, falling back to certifi's CA bundle on a verify failure.
@@ -75,7 +97,13 @@ def _urlopen(url: str, timeout: float = 60.0):
"""
try:
return urllib.request.urlopen(url, timeout=timeout)
except ssl.SSLCertVerificationError:
except (urllib.error.URLError, ssl.SSLCertVerificationError) as err:
# Only a certificate problem is worth a second attempt. "No route to
# host" and "connection refused" arrive as URLError too, and retrying
# those with a different CA list changes nothing except how long the
# operator waits for the real message.
if not _is_cert_failure(err):
raise
ctx = _https_context()
if ctx is None:
raise

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View File

@@ -4,7 +4,7 @@
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>Behavision</title>
<script type="module" crossorigin src="./assets/index-6LYbNlbD.js"></script>
<script type="module" crossorigin src="./assets/index-DSNu_2FW.js"></script>
<link rel="stylesheet" crossorigin href="./assets/index-DOJ2bRrM.css">
</head>
<body>

View File

@@ -14,6 +14,18 @@ import { MAKES, makeById } from '../../../../shared/cameraMakes.js'
// signed off through.
const BLANK = { id: '', host: '', port: 554, path: '', username: '', password: '', max_width: 1280 }
// "This computer's own camera" - a webcam or a built-in FaceTime camera.
//
// The engine has supported it since the first version (`webcam: 0` picks a
// capture index instead of building an RTSP URL) and no screen has ever
// offered it: another case of the API being able to do something the UI
// could not reach. It matters most for the thing it was missing from, which
// is showing the product to somebody. A laptop's own camera gives real
// recognition, of real faces, in the room, depending on no network at all -
// where pointing a demo machine at a camera in another building depends on
// two internet connections and a tunnel staying up while you talk.
const WEBCAM = 'webcam'
export default function Cameras() {
const { data, error, reload } = usePolled(() => api.cameras(), 8000)
const [editing, setEditing] = useState(null)
@@ -190,7 +202,9 @@ function useStreamURLs(cams) {
function CameraSheet({ cam, onClose, onSaved }) {
const isNew = !cam.id
const [f, setF] = useState({ ...BLANK, ...cam, password: '', path: cam.path || (isNew ? MAKES[0].path : '') })
const [make, setMake] = useState(isNew ? MAKES[0].id : 'manual')
const [make, setMake] = useState(isNew ? MAKES[0].id : (cam.webcam != null ? WEBCAM : 'manual'))
const [index, setIndex] = useState(cam.webcam != null ? String(cam.webcam) : '0')
const local = make === WEBCAM
const [test, setTest] = useState(null)
const [busy, setBusy] = useState(null)
const [error, setError] = useState(null)
@@ -199,6 +213,7 @@ function CameraSheet({ cam, onClose, onSaved }) {
// Only overwrite the path when the preset has one, so choosing "I know the
// path" does not wipe what the installer already typed.
function chooseMake(e) {
if (e.target.value === WEBCAM) { setMake(WEBCAM); setTest(null); return }
const m = makeById(e.target.value)
setMake(m.id)
setF(prev => ({ ...prev, path: m.path || prev.path }))
@@ -212,6 +227,14 @@ function CameraSheet({ cam, onClose, onSaved }) {
if (v === '' || v === null || v === undefined) continue
out[k] = (k === 'port' || k === 'max_width') ? Number(v) : v
}
if (local) {
// An address and a webcam index are alternatives, not extras: the
// engine's source() takes the webcam first, so leaving a half-typed
// host behind would make the saved camera describe two different
// things and only one of them would be used.
for (const k of ['host', 'path', 'username', 'password']) delete out[k]
out.webcam = Number(index) || 0
}
return out
}
@@ -259,7 +282,7 @@ function CameraSheet({ cam, onClose, onSaved }) {
<p className="lead">Three things from the camera: its address, its make, and its password. Test before you save — a wrong address is the most common mistake.</p>
{error && <div className="err"><Icon.Warning size={15} />{error}</div>}
{isNew && (
{isNew && !local && (
<section className="formsection">
<h4>Find it</h4>
{scan === null && (
@@ -301,15 +324,20 @@ function CameraSheet({ cam, onClose, onSaved }) {
<em className="hint">Short, no spaces. It names this camera everywhere and cannot be changed later.</em>
</label>
)}
<div className="fieldrow">
<label className="field"><span>Address</span>
<input value={f.host} onChange={set('host')} placeholder="192.168.1.20" inputMode="decimal" />
<em className="hint">On a label on the camera, or in its own app under “network”.</em>
</label>
<label className="field narrow"><span>Port</span>
<input value={f.port} onChange={set('port')} inputMode="numeric" />
</label>
</div>
{local
? <label className="field narrow"><span>Camera number</span>
<input value={index} onChange={e => { setIndex(e.target.value); setTest(null) }} inputMode="numeric" />
<em className="hint">0 is the built-in camera. Try 1 if a second one is plugged in.</em>
</label>
: <div className="fieldrow">
<label className="field"><span>Address</span>
<input value={f.host} onChange={set('host')} placeholder="192.168.1.20" inputMode="decimal" />
<em className="hint">On a label on the camera, or in its own app under “network”.</em>
</label>
<label className="field narrow"><span>Port</span>
<input value={f.port} onChange={set('port')} inputMode="numeric" />
</label>
</div>}
</section>
<section className="formsection">
@@ -317,27 +345,34 @@ function CameraSheet({ cam, onClose, onSaved }) {
<label className="field"><span>Make of camera</span>
<select value={make} onChange={chooseMake}>
{MAKES.map(m => <option key={m.id} value={m.id}>{m.label}</option>)}
<option value={WEBCAM}>This computer’s own camera</option>
</select>
{chosen.note && <em className="hint">{chosen.note}</em>}
</label>
<label className="field"><span>Stream path</span>
<input className="mono" value={f.path} onChange={set('path')} placeholder="/Streaming/Channels/101" />
<em className="hint">Filled in from the make. Change it only if the camera’s own app says something else.</em>
{local
? <em className="hint">Recognition runs on this computer’s built-in or plugged-in camera. Nothing on the network is involved.</em>
: chosen.note && <em className="hint">{chosen.note}</em>}
</label>
{!local && (
<label className="field"><span>Stream path</span>
<input className="mono" value={f.path} onChange={set('path')} placeholder="/Streaming/Channels/101" />
<em className="hint">Filled in from the make. Change it only if the camera’s own app says something else.</em>
</label>
)}
</section>
<section className="formsection">
<h4>Sign-in to the camera</h4>
<div className="fieldrow">
<label className="field"><span>Username</span>
<input name="rtsp-account" autoComplete="off" value={f.username} onChange={set('username')} placeholder="admin" />
</label>
<label className="field"><span>Password</span>
<input type="password" name="rtsp-secret" autoComplete="new-password" value={f.password}
onChange={set('password')} placeholder={cam.has_password ? '(unchanged)' : ''} />
</label>
</div>
</section>
{!local && (
<section className="formsection">
<h4>Sign-in to the camera</h4>
<div className="fieldrow">
<label className="field"><span>Username</span>
<input name="rtsp-account" autoComplete="off" value={f.username} onChange={set('username')} placeholder="admin" />
</label>
<label className="field"><span>Password</span>
<input type="password" name="rtsp-secret" autoComplete="new-password" value={f.password}
onChange={set('password')} placeholder={cam.has_password ? '(unchanged)' : ''} />
</label>
</div>
</section>
)}
{test && (
test.ok

View File

@@ -1,6 +1,6 @@
[project]
name = "behavision"
version = "1.1.0"
dynamic = ["version"]
description = "Production face recognition over RTSP"
requires-python = ">=3.10"
dependencies = [
@@ -22,7 +22,11 @@ dependencies = [
]
[project.optional-dependencies]
dev = ["pytest>=8.0"]
# httpx is test-only and never ships in the wheel. The HTTP tests begin with
# `pytest.importorskip("httpx")` so a bare checkout still runs - which is
# right, and has a cost worth knowing: without it the suite reports 239 passed
# and quietly SKIPS 32 API tests. `pip install -e .[dev]` is how to get them.
dev = ["pytest>=8.0", "httpx>=0.27"]
[tool.setuptools.packages.find]
include = ["behavision*"]
@@ -32,3 +36,9 @@ behavision = ["static/*"]
[tool.pytest.ini_options]
testpaths = ["tests"]
# One version for the package and the wheel. Two literals in two files is how
# they came to disagree (1.0.0 here, 1.1.0 there) and how neither tracked a
# release.
[tool.setuptools.dynamic]
version = {attr = "behavision.__version__"}

View File

@@ -50,7 +50,28 @@ step "2. Agent and setup tool"
step "3. Engine source and wheel"
# The wheel is built with the checkout's own interpreter; requires-python is a
# statement about the SHOP PC, which setup enforces when it finds Python there.
# Stamp the tag into the wheel version. Without this every release built
# behavision-1.1.0-py3-none-any.whl, and `pip install --upgrade` on a machine
# that already had 1.1.0 is a no-op - so an engine fix reached nobody who had
# ever run setup, while the Go binaries beside it updated normally. Nothing
# reported a version either, so there was no way to tell a shop PC three
# releases behind from a current one.
#
# v0.5.6-demo -> 0.5.6+demo, which is valid PEP 440: a local segment takes
# alphanumerics and dots, never hyphens.
TMPVER=$(mktemp)
PEP440=$(printf '%s' "${TAG#v}" | sed 's/-/+/; s/[^0-9A-Za-z.+]/./g')
trap 'git checkout -- behavision/__init__.py 2>/dev/null || true; rm -f "$TMPVER"' EXIT
# A temp file rather than `sed -i`, whose argument handling differs between
# BSD and GNU - this script is run from a Mac today and that is not a reason
# to plant a portability trap in a release path.
sed "s/^__version__ = .*/__version__ = \"$PEP440\"/" behavision/__init__.py > "$TMPVER"
cat "$TMPVER" > behavision/__init__.py
.venv/bin/python -m pip wheel --no-deps --ignore-requires-python -q -w "$STAGE/engine-src" . 2>&1 | grep -v "DEPRECATION\|WARNING: Ignoring" || true
git checkout -- behavision/__init__.py
rm -f "$TMPVER"
trap - EXIT
echo " engine wheel version $PEP440"
ls "$STAGE"/engine-src/behavision-*.whl >/dev/null || { echo "wheel was not built" >&2; exit 1; }
cp pyproject.toml requirements.txt "$STAGE/engine-src/"
mkdir -p "$STAGE/engine-src/config" && cp config/default.yaml "$STAGE/engine-src/config/"

View File

@@ -225,3 +225,42 @@ def test_an_inverted_per_camera_pair_is_rejected_not_stored(client):
tuning={"enroll_threshold": 0.8, "match_threshold": 0.5})
assert r.status_code == 400
assert c.get("/api/cameras").json() == []
# "This computer's own camera" — the demo case, and a capability the engine
# has always had with nothing able to reach it.
#
# It matters most for showing the product to somebody. A laptop's own camera
# gives real recognition, of real faces, in the room, depending on no network
# at all — where pointing a demo machine at a camera in another building
# depends on two internet connections and a tunnel staying up while you talk.
def test_this_computers_own_camera_can_be_added(client):
c, eng = client
r = c.post("/api/cameras", json={"id": "laptop", "webcam": 0})
assert r.status_code == 201, r.text
assert "laptop" in eng.workers
cam = next(x for x in c.get("/api/cameras").json() if x["id"] == "laptop")
# Returned, or the form cannot tell a webcam camera from a half-filled
# RTSP one when somebody opens it to edit.
assert cam["webcam"] == 0
assert eng.camera_store.get("laptop").source() == 0
def test_a_second_camera_index_is_kept(client):
"""0 is the built-in one; a plugged-in camera is usually 1. An index
silently coerced to 0 would open the wrong camera and look like the
setting had no effect."""
c, eng = client
assert c.post("/api/cameras", json={"id": "usb", "webcam": 1}).status_code == 201
assert eng.camera_store.get("usb").source() == 1
def test_a_webcam_camera_survives_a_reload(client, tmp_path):
"""It has to be on disk, not only in the running engine: a demo that
forgets its camera when the app restarts is worse than no demo."""
c, _ = client
c.post("/api/cameras", json={"id": "laptop", "webcam": 0})
again = CameraStore(tmp_path / "cameras.json").get("laptop")
assert again.source() == 0
assert again.safe_url() == "webcam:0"

View File

@@ -9,6 +9,7 @@ perfectly - numpy, onnxruntime, faiss, all of it - and then could not fetch a
"""
import io
import logging
import os
import ssl
import urllib.request
@@ -33,9 +34,18 @@ class _Resp(io.BytesIO):
def _verify_error():
return ssl.SSLCertVerificationError(
"""What urlopen ACTUALLY raises, which is not what it looks like.
urllib catches ssl.SSLCertVerificationError and re-raises
urllib.error.URLError(err), carrying the original on `.reason`. The first
version of this test raised the bare SSL error - a shape real urllib never
produces - so it passed against a fallback that could never fire, and the
fix shipped and failed on the machine it was written for with the exact
traceback it was meant to prevent.
"""
return urllib.error.URLError(ssl.SSLCertVerificationError(
"[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: "
"unable to get local issuer certificate")
"unable to get local issuer certificate"))
def test_a_machine_with_no_trust_store_falls_back_to_the_bundled_one(monkeypatch):
@@ -74,12 +84,22 @@ def test_a_working_trust_store_is_used_as_is(monkeypatch):
def test_a_real_network_failure_is_not_disguised_as_a_certificate_problem(monkeypatch):
"""A URLError is not automatically a certificate problem.
"No route to host" and "connection refused" arrive as URLError too, and
retrying those with a different CA list changes nothing except how long
the operator waits for the real message.
"""
calls = []
def fake(url, timeout=None, context=None):
calls.append(context)
raise urllib.error.URLError("no route to host")
monkeypatch.setattr(urllib.request, "urlopen", fake)
with pytest.raises(urllib.error.URLError):
model_assets._urlopen("https://example.invalid/m.onnx")
assert calls == [None], f"a dead network was retried as a CA problem: {calls}"
def test_progress_lines_survive_the_rewrite(tmp_path, monkeypatch, caplog):
@@ -103,3 +123,42 @@ def test_progress_lines_survive_the_rewrite(tmp_path, monkeypatch, caplog):
pct = [ln for ln in lines if ln.startswith("download: face detector ")]
assert pct, f"no progress lines at all: {lines}"
assert "download: face detector 100%" in pct, pct
@pytest.mark.skipif(not os.environ.get("BEHAVISION_NETWORK_TESTS"),
reason="set BEHAVISION_NETWORK_TESTS=1 to reach github.com")
def test_against_a_real_machine_with_no_trust_store(tmp_path, monkeypatch):
"""The whole thing, over the real network, with the real failure.
Every stub above is a statement about what I believe urllib does, and the
first version of this file proved how much that is worth. Python's default
context honours SSL_CERT_FILE, so pointing it at an EMPTY file reproduces
a python.org macOS build exactly: a default context that trusts nobody.
certifi is loaded by path and is unaffected by the variable, so if the
fallback works here it works there.
"""
empty = tmp_path / "no-cas.pem"
empty.write_text("")
monkeypatch.setenv("SSL_CERT_FILE", str(empty))
monkeypatch.setenv("SSL_CERT_DIR", str(tmp_path / "nothing"))
# Not every interpreter can be put into that state, and pretending
# otherwise would turn this into a test that passes by not running. A
# Python linked against LibreSSL - which is what macOS Command Line Tools
# ships - reads the system keychain and ignores SSL_CERT_FILE entirely:
# measured, 128 CAs loaded with the variable pointing at an empty file. A
# python.org build links OpenSSL and honours it: 0 CAs, which is the
# condition being reproduced.
if ssl.create_default_context().get_ca_certs():
pytest.skip("this interpreter ignores SSL_CERT_FILE (%s), so the "
"no-trust-store condition cannot be reproduced here"
% ssl.OPENSSL_VERSION)
# The premise: the default context really is broken in this process.
with pytest.raises(urllib.error.URLError) as caught:
urllib.request.urlopen(model_assets.YUNET_URL, timeout=30)
assert model_assets._is_cert_failure(caught.value), caught.value
dest = tmp_path / "yunet.onnx"
model_assets._fetch(model_assets.YUNET_URL, dest, "face detector")
assert dest.stat().st_size > 200_000

View File

@@ -0,0 +1,81 @@
"""Why a camera "will not connect", when the real answer is that the computer
asking is in the wrong building.
Asked directly by the owner, about his own cameras, from his phone's
connection: *"when i connect from my mobile internet the cameras wont connect,
why is that"*. The answer was `cannot reach 192.168.1.121:554 - Operation
timed out`, which reads as a broken camera and sends somebody to re-type an
address and a password that were always correct.
"""
import pytest
from behavision import capture
@pytest.fixture
def on(monkeypatch):
def _set(ip):
monkeypatch.setattr(capture, "_local_ipv4", lambda: ip)
return _set
def test_a_different_network_says_so_and_says_what_to_do(on):
on("10.11.12.13") # a phone's tethered network
hint = capture._wrong_network_hint("192.168.1.121")
assert "not the camera's network" in hint
# The action, not just the diagnosis: no setting on this screen fixes it.
assert "computer in the shop" in hint
def test_the_same_network_sends_you_to_the_camera_instead(on):
on("192.168.1.120")
hint = capture._wrong_network_hint("192.168.1.121")
assert "on that network" in hint
assert "powered on" in hint
# Two states that need opposite actions must not share a sentence.
assert "not the camera's network" not in hint
@pytest.mark.parametrize("host", [
"8.8.8.8", # plainly routable
"203.0.113.9", # TEST-NET-3: `is_private` calls this private, and it is
# not a LAN address - which is why the check spells out
# the RFC1918 blocks instead of asking is_private
"100.64.0.5", # carrier-grade NAT, what a mobile network hands out
])
def test_an_address_that_is_not_a_lan_address_gets_no_hint(on, host):
"""A routable address unreachable from here is an ordinary network fault,
and inventing a story about private networks would be wrong."""
on("192.168.1.120")
assert capture._wrong_network_hint(host) == ""
def test_a_name_gets_no_hint(on):
"""Nothing can be concluded about `camera.local` from the string, and a
guess here is a confident wrong answer in the place people look first."""
on("192.168.1.120")
assert capture._wrong_network_hint("camera.local") == ""
def test_loopback_gets_no_hint(on):
on("192.168.1.120")
assert capture._wrong_network_hint("127.0.0.1") == ""
def test_with_no_network_at_all_it_still_names_the_cause(on):
"""A machine with no route cannot say which network it is on, and must not
pretend: the private-address fact is still true and still the reason."""
on("")
hint = capture._wrong_network_hint("192.168.1.121")
assert "private address" in hint
assert "this computer is on" not in hint
def test_the_hint_reaches_the_message_a_person_reads():
"""The whole point is the sentence on the screen, not a helper nobody
calls. Port 1 on a private address refuses or times out immediately."""
ok, msg = capture._tcp_reachable("rtsp://192.168.1.121:1/ch0", 1.0)
assert not ok
assert "192.168.1.121" in msg
# Whichever branch the OS takes, the explanation travels with it.
assert "network" in msg