Brand case decides whether the catalogue-intelligence service answers at all.
Measured 30 Sep 2026:
/nutrition/Balaji/balaji_..._135g -> health_score 65.3, 545 kcal
/nutrition/balaji/... (our spelling) -> every field null
The brand list resolves ours to theirs, and /brands has slowed to 0.2-2.3s,
which exceeded the 3s client timeout on a cold start. The fallback then asked
under our own spelling, received a well-formed empty record, and cached it as
"no nutrition" for six hours -- so one slow moment silently removed nutrition
and health scores from every product of every brand, looking exactly like data
the agent team had not supplied.
Two changes:
- a result reached without a resolved brand is no longer cached, so the next
request retries rather than inheriting a wrong answer for six hours. A
genuine miss on a resolved brand is still cached, which is the case that
matters for traffic.
- the brand list is warmed in the background at startup, so no shopper is
ever in the path of that call.
Also logs which state the feature is in at boot, the way mail does. With
NUTRITION_BASE unset the endpoint simply omits `nutrition` and `healthscore`,
which is indistinguishable from an unscored product -- this deploy went out
without the variable set and had to be diagnosed by probing the API from
outside.
scratch/nutritionlive prints the exact response for any product by running this
code against the live product row and the live service.
NUTRITION_BASE=https://mcp.nearle.ai.in/api must be set in the deployment
environment. Unset, nothing changes and no product carries either key.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 2: the loop and the model gateway. The composer in the console has
said "Not connected yet" since it was built, because there was no
assistant endpoint anywhere. There is one now.
- utils/chat.go the gateway, a sibling of embedding.go: one small
interface, a provider switch, the shared postJSON, no
framework. Agents name a TIER (fast/balanced/deep) and
config maps tier to model, so changing provider does not
touch an agent.
- services/assistantService.go one loop for every agent. An agent is a
name, a tier, a prompt and an allow-list — data, not a
class — so a sixth is config rather than a subclass.
- the endpoint under /v1/web, inheriting middleware.WebAuth along with
every other console route. The assistant reads the same data the console
does and must read it as the same person.
What the model does not get to decide:
whose data the caller is built from the verified session in the
controller, never from the request body — there is no
tenant field to fill in. A test scripts the model calling
a tool with {"tenantid": 916} and asserts it ran for 1147.
which tools the registry enforces the agent's allow-list; a test
scripts a call to a tool the agent lacks and asserts the
handler never ran.
when to stop steps and tool calls are counted here. A model that keeps
calling tools is stopped by arithmetic, not by being
asked nicely.
Two quiet failures have tests of their own. A finish_reason of "length"
means the provider cut the reply off mid-sentence, which reads exactly
like a complete answer unless it is flagged. And a truncated tool result
reaches the model in words it will repeat — otherwise it describes a
capped list and an empty one identically.
A refused tool goes back as a message, not an error: a model told "that
tool needs a tenant" can explain it, where a model handed nothing says
"something went wrong".
Optional, like the embedder. Without ASSISTANT_PROVIDER the endpoint
answers "not switched on here", the composer stays disabled, and the tools
still work — they are ordinary Go functions, and only turning a sentence
into a tool call needs a model.
14 tests, against a scripted model rather than a live provider: these are
about what the loop refuses to let a model do, and that has to hold for
any model, including one behaving badly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 1 of Nearle Buddy: an agent names a tool, and the registry decides
whether that is allowed, whether the arguments make sense, who is asking,
and what gets recorded — then runs a handler a person wrote and tested.
No agent gets raw table access. The usual argument for tools over
generated SQL is safety; here there is a harder one. The fields on this
backend do not mean what their names say, and it is measured:
orders.deliverystatus is an empty string on all 181 rows of tenant 1147,
orders.orderstatus never carries the six middle delivery stages,
deliveries.ridername holds statuses as often as names, deliverytype is
empty on every row in production. A model writing SQL gets each of those
wrong with no error — it reports a cancel rate from a column of empty
strings and nobody can tell. A model calling a tool cannot, because the
correction lives in the handler beside the measurement that justified it.
Call does five things in order: find the tool, check the agent's
allow-list, validate arguments, confirm the caller is scoped to
something, run the handler — writing exactly one audit row whatever
happens, refusals included. A trail of successes answers "did anything
try to read another tenant?" with silence, which reads the same as no.
The model has no say in whose data is read. stuck_orders has no tenantid
field on its schema — absent, not rejected — and the tenant comes from
the session claims added in the previous commit. Arguments the tool did
not declare are dropped rather than passed on, so a model sending a
`where` clause gets it discarded.
stuck_orders: deliveries a rider was given and has not accepted, ten
minutes for a look, twenty-five for somebody now. Derived from assigntime
and orderstatus, so it does not depend on anyone having been watching.
Carries the wait in minutes, what to do, where to check it, and what it
covered. A capped answer says so — an empty result and a truncated one
look identical to a model and it will call both "none".
The audit sink writes to the log for now; a database sink is phase 8.
Nothing calls the registry yet: the loop and the model gateway are phase 2.
37 tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
main had moved on with retrieval work validated against real queries —
minTokenHits (the word match needs two thirds of the label, not all of
it), separator folding so "Parle G"/"Parle-G"/"ParleG" all reach Parle-G,
the floor at 0.50 after "Paracetamol" came back as "Paneer Makhni 500ml"
at 0.304, and ties broken on cosine distance instead of name. All of that
is kept exactly as it was.
The conflict was in textScore: this branch replaced the substring rule
with a coverage formula to stop a bare brand name resolving to one
arbitrary product. That is the wrong half to change. The substring rule
scores every product of a brand 0.95 IDENTICALLY, and that tie is not the
bug — it is the signal. isAmbiguous reads it, so the branch's coverage
rewrite is dropped and the ambiguity layer alone does the work:
"britannia" → all 258 rows tie at 0.95 → ambiguous: true + candidates
"Parle G" → folding and the single-character token still land it
a real name → runner-up far behind → match, unchanged
Dropped with it: scanSpecificEnough, the per-hit text score, and the
proportional confirmation bonus — the flat +0.10 is back. Simpler, and it
leaves main's tuning untouched.
TestTextScoreRewardsSpecificityNotJustOverlap tested the removed formula
and is replaced by TestABrandNameScoresItsProductsIdentically, which
guards the tie itself: a formula that broke it on name length or word
count would bring the bug back.
Docs carry both rationales, and now say plainly that confidence stays
high on the ambiguous path — gate on `ambiguous`, never on `confidence`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`"britannia"` is a substring of all 258 Britannia product names, and
textScore returned 0.95 for any product whose name contained the label.
So every one of them tied, the tie broke alphabetically, and the customer
was shown one arbitrary biscuit with "confidence": 0.95 and a price. Lens
hands back a bare wordmark often — it is usually the biggest thing printed
on a packet — so this was the common case, not an edge one. Found via the
example request in the mobile team's own proposal.
Scoring now asks both questions. A hit carries `score` (ranks) and `text`
(how specifically the label names THIS product: the harmonic mean of how
much of the label the product explains and how much of the product's name
the label explains, pack sizes dropped from both sides). A brand name
scores its products ~0.33 equally instead of 0.95 arbitrarily. The
"vector and text agree" bonus is now proportional to the text score, so a
weak match can no longer inflate a whole brand.
isAmbiguous reads that: the leader is a guess if anything is level with it
(margin) or if the label names no one product (specificity), and then the
response carries `ambiguous: true` with `candidates` — distinct products,
not pack sizes, at most ten, each marked with whether one of the
customer's stores has it in stock, available ones first. `match` is nil
and `stores` empty on that path: no price for a product nobody chose.
Erring towards asking is deliberate — a tap versus the wrong biscuit.
To act on a pick, /lookup now accepts `brand` + `catalogueid` instead of a
label and skips recognition entirely (also serves deep links and re-order).
New: ScanRepository.CatalogueRef, resolving via the brand tables discovered
from information_schema, never a name built from the request.
Also: scratch/cataloguedims now reports every vector column, not just
`embedding` — which is how we learned the catalogue also carries
img_vector(1024), filled on 1885 of 2124 rows. SCAN_TO_ORDER.md records
why that column stays unread for now and what would change it, alongside
why the app is not asked to compute vectors on the phone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
POST /v1/mob/scan/lookup label + customer → catalogue match, sizes, and
every registered store that sells it with live
stock, in-stock first / nearest first, one
recommended
POST /v1/mob/scan/confirm chosen store + size + qty → re-read the ledger;
ok, or the next-nearest store with enough of the
same product
GET /v1/mob/scan/stores registered stores nearest first
Recognition is pgvector cosine search over every brand_* table (each
with its own index, merged) plus a word match that settles near-ties
and works alone when no model is configured. The embedder is chosen by
EMBEDDING_PROVIDER (OpenAI-compatible or Gemini) and must be the model
that indexed the catalogue: verified 2026-09-15 as all-MiniLM-L6-v2 over
search_query, served by the cluster's Ollama as `all-minilm`; the first
search refuses a width mismatch by name.
Customer, stores and catalogue are read concurrently under a 5 s cap; a
slow model degrades to a text answer. Vectors and ranked hits are cached
in Redis and in-process; live stock never is. Availability uses the same
rules as the customer catalogue (approve, publishedat, ledger balance,
outlet price else retail). No stock reservation: confirm re-reads.
scratch/cataloguedims reports the catalogue's embedding width and fill.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
GetUserByAuthname / GetUserByContactNo / GetUserLogin discarded the
Scan error, so a database that could not answer — down, pool exhausted,
or booted without its config (2026-07-20) — came back as uid 0 and every
user was told their email was wrong.
One lookup, GetUserLogin, now returns an error; sql.ErrNoRows is "not
found" and anything else reaches the service, which answers 500 "Login
is temporarily unavailable" and logs the cause. 409 "Invalid Email" is
unchanged for a genuine no-match: the console reads that exact shape as
"not registered". NULL password/role columns scan through sql.Null* so
they do not become 500s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>