Image capturing Flow updates

This commit is contained in:
sriram
2026-09-24 17:18:38 +05:30
parent f933ea10a1
commit fe208f4715
14 changed files with 1877 additions and 6 deletions

View File

@@ -0,0 +1,219 @@
# Capture-to-Catalog: workflow
**What this covers:** what happens when a colleague photographs a product,
including one the catalogue does not have yet (for example *Godrej Fab*,
Detergents & Fabric Care, missing from `brand_godrej`).
**The requirement:** the colleague must never be told "No such product found".
The system reads the label, researches the product, adds a proper record to the
brand table, and returns the details.
**Status (2026-09-24):** built and tested. It is **off by default**
(`ENABLE_CAPTURE_DISCOVERY=false`) until the rollout steps below are done.
---
## 1. Before and after
| | Before | After (flag on) |
|---|---|---|
| Product in catalogue | Returned by image or label match | Same, unchanged |
| Product NOT in catalogue | Returned the *closest other* products (e.g. a rival detergent) marked unconfirmed. Nothing was added. | Instant card from the label + background job that adds the product. The colleague gets the full record when it's ready. |
| Label unreadable | Best-effort list, unconfirmed | A clear ask: "retake closer, or type the product name" |
| Brand never seen before | Best-effort list | A clear ask; **no brand table is created from a photo** |
**Was it a cache problem? No.** Identify results are never cached. The
real reasons a product could stay invisible are fixed or explained in §6.
---
## 2. The workflow, step by step
```
Colleague takes photo (Nearle app)
│
▼
POST /api/search/identify
│
┌────────┴──────────────────────────────────────────────┐
│ STEP 1 Identify (read-only, as before) │
│ a. Image match: photo vector vs img_vector (≥ 0.70) │
│ b. Label match: OCR text → catalogue text search │
└────────┬──────────────────────────────────────────────┘
│
confirmed? ── yes ──► return the product. DONE.
│ no
▼
┌───────────────────────────────────────────────────────┐
│ STEP 2 Read the label (capture_discovery.parse_label)│
│ brand "Godrej" (from the brand word) │
│ product "Godrej Fab Detergent Powder" │
│ pack size "1kg" │
│ category "Detergents & Fabric Care" │
│ HSN / GST 3402 / 18% │
└────────┬──────────────────────────────────────────────┘
│
unreadable, or brand not in catalogue ──► needs_input
│ ("type the brand / name")
▼
┌───────────────────────────────────────────────────────┐
│ STEP 3 Duplicate check │
│ same image_id already stored ──► exists: return it │
│ same product already queued ──► that job's id │
│ device over hourly limit ──► busy + card │
└────────┬──────────────────────────────────────────────┘
▼
┌───────────────────────────────────────────────────────┐
│ STEP 4 Reply at once (colleague is not kept waiting) │
│ matched_by "discovery_pending" │
│ provisional card (from step 2) + discovery_job_id │
└────────┬──────────────────────────────────────────────┘
│ background worker, one job at a time
▼
┌───────────────────────────────────────────────────────┐
│ STEP 5 The existing 11-stage pipeline, on one row │
│ 1 brand + FSSAI 2 LLM description │
│ 3 title/category 4 pack size │
│ 5 price band 6 web images (identity-checked) │
│ 7 SKU 8-9 barcode, HSN/GST, content │
│ 10 validation gate 11 embed + write to brand_godrej │
└────────┬──────────────────────────────────────────────┘
▼
┌───────────────────────────────────────────────────────┐
│ STEP 6 Mark the new row (only if it was INSERTED) │
│ live retail check: is this exact pack sold online? │
│ validation_status = needs_review │
│ field_sources.capture = origin, job, retail verdict │
│ no web image? → colleague's photo becomes the image │
└────────┬──────────────────────────────────────────────┘
▼
App polls GET /api/search/identify/jobs/{job_id}
→ status "done" + the full stored product
```
### Timing the colleague sees
- **~1–3 s:** the identify response with the provisional card
(brand, name, size, category, HSN) read from the label.
- **About a minute or more:** the full record: description, images, price
band, SKU and tax fields. It is slower when jobs queue up, because one runs at
a time.
- **Next photo of the same product:** found straight away through the label
match. It also matches on the photo itself once the image vector exists.
---
## 3. What the app receives
**Identify response, new fields** (additive; existing fields unchanged):
| Field | Meaning |
|---|---|
| `discovery_status` | `pending` / `exists` / `needs_input` / `busy`; `null` when the answer was already confirmed or the feature is off |
| `discovery_job_id` | Id to poll (only for `pending`) |
| `discovery_message` | One line to show the colleague |
| `provisional` | The card read from the label, including `visible_in_search` |
A **confirmed** answer is now: `matched_by` = `text`, `label_exact`, or
`image_vector` with no `fallback_reason`.
**Job poll:** `GET /api/search/identify/jobs/{id}` returns
`queued → running → done | rejected | failed | interrupted`. When the status is
`done`, `product` holds the stored catalogue card. `retail_presence` shows
whether a retailer lists that exact pack.
**Admin review list:** `GET /api/admin/captures` (admin only) shows recent
capture jobs, newest first.
Full contract: `backend/docs/IMAGE_SEARCH_API.md`, section *"When the product
is not in the catalogue"*.
---
## 4. Decisions taken
| # | Decision | Why |
|---|---|---|
| D1 | **Instant card + background job**, not one long request | The full pipeline takes minutes. A mobile request would time out. |
| D2 | New rows are **visible immediately, marked `needs_review`** | The photo proves the product exists, but OCR can misread a letter or the pack size. A human confirms later. |
| D3 | The **colleague's photo is used as the image only when the web finds none** | Clean retailer images look better. A product with no image at all is worse than a shelf photo. |
| D4 | **Only brands already in the catalogue** | An OCR misread must never create a junk brand table. Brands outside `ACTIVE_BRANDS` are still written, but stay hidden from search until switched on. |
**Safety rules built in:**
- A capture never overwrites or downgrades an existing row. The review mark is
applied only to a row the job itself inserted. A `rejected` verdict is never
lifted.
- A product-line word alone ("Fab") is never treated as a brand, because in the
registry "Fab" is Parle's biscuit.
- Per-device hourly limit (`CAPTURE_MAX_PER_CLIENT_PER_HOUR`, default 20) and
a bounded queue (`CAPTURE_QUEUE_MAX`, default 8), since the route is public.
- If discovery errors, identify still answers. It falls back to `needs_input`
and never fails with a 500.
---
## 5. Where the code is
| Piece | File |
|---|---|
| Label parsing, duplicate check, jobs, worker, row marking | `backend/app/services/capture_discovery.py` (new) |
| Hook into identify, job/photo/admin routes | `backend/app/api/routers/search.py` |
| Response models | `backend/app/api/schemas.py` (`IdentifyOut` fields, `ProvisionalProductOut`, `CaptureJobOut`) |
| Settings | `backend/app/infrastructure/settings.py` (`ENABLE_CAPTURE_DISCOVERY`, `CAPTURE_*`) and `.env.example` |
| Brand fix ("Godrej Fab" → godrej) | `backend/app/services/brand_registry.py` (`_known_leading_brand`) |
| Category fix (detergents → "Detergents & Fabric Care") | `backend/app/services/category_registry.py` |
| Tests | `tests/test_capture_discovery.py`, `tests/test_brand_registry.py`, `tests/test_category_registry.py` |
The 11-stage pipeline itself (`store_catalog_pipeline.run_pipeline`) is
**reused unchanged**.
---
## 6. Things that hide a product, and their status
| Issue | Effect | Status |
|---|---|---|
| "Godrej Fab" did not resolve to Godrej | Would have created a separate `brand_godrej_fab` table | **Fixed.** A known brand followed by a product line now joins that brand. |
| Detergents were categorised as "Household Cleaning" | Every detergent got its tax code flagged for review | **Fixed.** They now use "Detergents & Fabric Care" (HSN 3402, 18%, no flag). |
| `ACTIVE_BRANDS` excludes Godrej | Godrej Fab is stored but hidden from search | **Needs a decision:** add Godrej to `ACTIVE_BRANDS` (or leave it blank) and restart |
| Image vector is filled after the write | The next photo matches on the label first, not the image | By design. It catches up within seconds to minutes. |
---
## 7. Known limits
1. **Label reading is rule-based.** It needs a known brand word on the label.
Heavily stylised or partial labels get `needs_input`, and the colleague types
the name (sent as `text` / `brand`). An LLM fallback for messy labels is a
possible later addition.
2. **One vector per product.** When the web finds an image, the image vector
is built from that image, not the colleague's photo. Re-captures are then
matched by the label, which works because the row now exists. Matching on the
colleague's own photo in every case would need a second vector column.
3. **Non-food has no Open Food Facts source.** For detergents, the live retail
check is the only outside confirmation. That is why rows stay `needs_review`.
4. **The photo is shown as the image only when `CAPTURE_PUBLIC_BASE_URL` is set**,
because image URLs must be absolute. It is not uploaded to S3.
5. **Only `POST /search/identify`** (photo upload) triggers discovery. The
vector-only route `/search/image-vector` has no photo and does not.
---
## 8. Rollout steps
1. **Check production for split brand tables:**
`SELECT table_name FROM information_schema.tables WHERE table_name LIKE 'brand\_%';`
Look for any table named after a known brand plus a product line. This is
the one thing the brand fix could re-route.
2. Decide `ACTIVE_BRANDS`: add Godrej and any other brands colleagues will capture.
3. Set `CAPTURE_PUBLIC_BASE_URL` if the shelf photo should be used as the image
for products the web has no image for.
4. **Smoke test against the local database, not production.** `backend/.env`
currently points at **production**. A throwaway brand cannot be used, because
D4 refuses unknown brands on purpose. So a smoke test always writes a real
`needs_review` row into a real brand table: run it locally, or delete the row
afterwards (its `image_id` is in the job record).
5. Turn it on: `ENABLE_CAPTURE_DISCOVERY=true`, then restart the API.
6. Tell the Nearle app team about the new fields (§3). The app should show
`provisional` + `discovery_message` and poll the job.
7. Review new rows regularly: `GET /api/admin/captures`, or query
`validation_status = 'needs_review'` with `field_sources ? 'capture'`.

View File

@@ -0,0 +1,239 @@
# After deployment: what changes
**Scope:** the changes made on 2026-09-24:
- brand-alias fix
- detergent category fix
- capture-to-catalog flow
How the flow works is described in `CAPTURE_TO_CATALOG_WORKFLOW.md`. This
document covers what the running system does differently once these changes
are deployed.
---
## 1. Summary
The changes fall into two groups:
| Group | When it takes effect | Needs any action? |
|---|---|---|
| **A. Live on deploy.** Category detection, brand-name resolution, extra response fields, new routes. | As soon as the new build starts | No. Verify with the checks in §6. |
| **B. Dormant.** Capture-to-catalog (adding products from photos). | Only after `ENABLE_CAPTURE_DISCOVERY=true` and a restart | Yes. Decisions and steps in §4. |
**What the deploy does not change:**
- no database migration or new columns;
- no new container or Python dependency (rapidocr is already in the image);
- no change to existing rows.
With the flag off, no route writes anything new.
---
## 2. Live on deploy (flag-independent)
### 2.1 Detergent searches return detergents
**What changed:**
- `detergent`, `laundry`, `washing powder`, `detergent powder`, `detergent bar`,
`fabric wash` and `fabric care` now map to **"Detergents & Fabric Care"**.
- They used to map to "Household Cleaning".
**Effect on search and chat** (`catalog_search`, `rag_service`):
- A query like "detergent" or "Surf Excel detergent" is limited to the
detected category.
- **Before:** it was limited to "Household Cleaning". The last production
snapshot has about **38 rows** stored as "Detergents & Fabric Care" and
**2** as "Household Cleaning", so the real detergents were **filtered out**.
- **After:** those detergent rows are returned.
- "Household Cleaning" keeps dishwash, floor cleaner, handwash and cleaner.
**Effect on new uploads:**
- A detergent row gets category "Detergents & Fabric Care" and HSN **3402 at
18%** with **`hsn_gst_needs_review = false`**. Before, it was flagged for review.
- Stored rows keep their category. Nothing is recategorised.
**Risk:** if either of the 2 "Household Cleaning" rows is actually a detergent,
a "detergent" search no longer returns it. It is still found by its name or brand.
### 2.2 "Brand + product line" joins the parent brand
**What changed:** `resolve_parent_brand` now has one more rule, used only
when every older rule found nothing. A name that starts with a known brand
resolves to that brand:
| Input | Before | After |
|---|---|---|
| Godrej Fab | own table `brand_godrej_fab` | `brand_godrej` |
| Amul Taaza | own table `brand_amul_taaza` | `brand_amul` |
| Hindustan Unilever Rin | own table | `brand_hindustan_unilever` |
| Fab Detergent Powder | unchanged | unchanged (a line name alone is never treated as a brand) |
Anything that resolved before resolves the same way. All seed catalogs are
pinned by tests.
**Where it applies:**
- uploads (stage 1 of the pipeline);
- per-brand reads such as `/api/brands/{brand}/products`;
- the S3 folder and the seed-file choice.
**Risk (not verified, because production could not be read from here):**
production may already have a table created under such a name, for example a
`brand_amul_taaza` from an earlier upload. If so, requests for that brand name
now read and write the parent table instead, and the old table becomes
orphaned. Its rows are not lost, only unreachable by that name. Check this
before deploying (§6, check 1).
### 2.3 Identify response has four extra keys
`POST /api/search/identify` (and `/search/image-vector` with
`text_fallback`) now always include these keys:
`discovery_status`, `discovery_job_id`, `discovery_message`, `provisional`
With the flag off they are all `null`. No existing key changed. The Nearle
app is unaffected unless it rejects unknown JSON keys.
### 2.4 New routes (visible in the OpenAPI docs)
| Route | With the flag off |
|---|---|
| `GET /api/search/identify/jobs/{job_id}` | always 404 (no jobs exist) |
| `GET /api/admin/captures` | admin only; returns an empty list |
| `GET /api/search/captures/{name}` | always 404 (no photos exist); hidden from the docs |
---
## 3. Which environment file the server uses
- **The image:** `backend/Dockerfile` copies `.env.production` into the image
as `.env`.
- **Compose:** `docker-compose.yml` also loads `./backend/.env` from the
server's checkout into the backend container. Values from compose win.
| Setting | `.env.production` | local `backend/.env` |
|---|---|---|
| `ENABLE_CAPTURE_DISCOVERY` | not set, so **off** | not set, so **off** |
| `ACTIVE_BRANDS` | not set, so **all brands visible** | `Amul, Cadbury, Hindustan Unilever, Own Products` |
The capture flow is off in production however it is deployed.
Whether Godrej products appear in search depends on which file the server
actually loads. Confirm with `GET /api/brands` after deploy (§6, check 4).
---
## 4. After `ENABLE_CAPTURE_DISCOVERY=true`
### 4.1 What happens on each request
Each identify request that **cannot be confirmed** now also does the following:
1. Reads the label into brand, product name, pack size and category. This adds
a few milliseconds.
2. Chooses what to return:
- known brand, new product: queues a job and returns `pending` with a
provisional card. The unconfirmed lookalikes are no longer in `results`.
- product already stored: returns it as a confirmed match (`label_exact`).
- unreadable label or unknown brand: returns `needs_input` with a message.
- rate limit or full queue: returns `busy` with the card.
Confirmed identifies behave exactly as before.
### 4.2 What a job does in the background
- **One worker thread** in the backend process runs one job at a time.
- Each job runs the full 11-stage pipeline for one product, which makes
outbound calls:
- an Ollama description, up to 120 s;
- web image search, several providers, validated one by one;
- a live retail-presence search;
- an SKU lookup if `ENABLE_SKU_WEB_LOOKUP` is on.
- **Expect about a minute or more per job.** Later jobs wait in the queue
(up to `CAPTURE_QUEUE_MAX = 8`).
- It then writes **one row** to the brand table:
- `validation_status = needs_review`;
- `field_sources.capture` records the origin, job id, label text and retail verdict.
- The image vector is computed by the existing background worker. If the
colleague's photo became the product image, its vector is stored straight away.
### 4.3 Resources
| Resource | Impact |
|---|---|
| CPU / RAM | One pipeline run at a time, inside the backend's existing 2.5 GB limit. It can run alongside an admin batch upload; both use Ollama, so each gets slower. |
| Network | Outbound calls to image search, retail search and Ollama for each job |
| Disk | `/app/data/captures/`, on the `catalog_rag_backend_data` volume, so it survives redeploys. One photo (roughly 0.5–3 MB) plus a small JSON record per job. **There is no automatic cleanup yet.** |
| Database | One row per new product, plus two small targeted updates |
### 4.4 Restarts
- A queued or running job does not survive a restart. When it is next polled
it reports `interrupted`, and the colleague captures again.
- Rows that were already written stay in the table.
- The per-device rate-limit counters reset.
### 4.5 Safety rules
- **Only known brands.** A registry parent or an existing brand table.
A photo can never create a new brand table.
- **Never overwrites.** The `needs_review` mark and provenance are applied only
to a row the job itself inserted. An existing row is left exactly as it was,
and a `rejected` verdict is never lifted.
- **Public route limits.** Each device (by IP) can start 20 jobs per hour
(`CAPTURE_MAX_PER_CLIENT_PER_HOUR`), and the queue is bounded. Behind a
proxy, several devices can share one IP.
- **Fails safe.** Any error inside discovery turns into `needs_input`.
Identify never returns a 500 because of it.
---
## 5. How to enable it
1. **Check production for tables split by product line** (see §6, check 1).
2. **Decide `ACTIVE_BRANDS`** for the server. Blank means every brand is
visible. Otherwise list every brand colleagues will photograph, Godrej included.
3. **Optional:** set `CAPTURE_PUBLIC_BASE_URL=https://api.<domain>`. This lets
the shelf photo be used as the image for products the web has no image for.
4. Set `ENABLE_CAPTURE_DISCOVERY=true` in the env file the server actually loads
(§3), then restart the backend.
5. **Tell the Nearle app team** about the new response fields and the job poll
(`IMAGE_SEARCH_API.md`, "When the product is not in the catalogue").
6. **Smoke test with one real product** and delete the row afterwards if it
was only a test. A made-up brand can't be used, because the brand gate
refuses it on purpose.
---
## 6. Post-deploy checks
| # | Check | Expected |
|---|---|---|
| 1 | *(before deploy)* `SELECT table_name FROM information_schema.tables WHERE table_name LIKE 'brand\_%';` | No table named after a known brand plus a product line (e.g. `brand_amul_taaza`). If one exists, decide whether to merge it first. |
| 2 | `GET /api/health` | Healthy. `ocr` is available (the capture flow needs server OCR when the app sends no `text`). |
| 3 | `GET /api/search?q=detergent` | Detergent products are returned (category "Detergents & Fabric Care"). |
| 4 | `GET /api/brands` | The brand list matches your `ACTIVE_BRANDS` decision. |
| 5 | `POST /api/search/identify` with any photo | The same answer as before, plus four `null` discovery keys. |
| 6 | `GET /api/admin/captures` (admin token) | `[]` |
| 7 | *(after enabling)* Photograph a product that is not in the catalogue | `discovery_status: "pending"`, then the job reaches `done` and `product.validation_status` is `needs_review`. |
---
## 7. Rollback
| What | How | What remains |
|---|---|---|
| Stop adding products | `ENABLE_CAPTURE_DISCOVERY=false`, then restart | Rows already added, which can be found by the query below |
| Remove the whole change | Redeploy the previous build | The same rows. Nothing needs to be migrated back. |
| Remove the rows it added | `DELETE FROM brand_<x> WHERE field_sources ? 'capture' AND validation_status = 'needs_review';` (review first with a `SELECT`) | Photos under `/app/data/captures/`, which can be deleted |
Reverting the category or brand fix affects **new** uploads only. Rows
written while the fix was live keep the values they were given.
---
## 8. Open items
1. **Photo and job retention.** There is no cleanup job yet, so photos build
up on the data volume. A purge after N days is worth adding before heavy use.
2. **The photo route is public.** It is reachable by anyone who has the
32-character job id. The ids can't be guessed, but anyone with a job id can
view that photo.
3. **One image vector per product.** When the web supplies the image, the
colleague's photo is not used for image matching. Repeat photos of that
product are matched by reading the label instead.
4. **Rule-based label reading.** It needs a known brand word printed on the
pack. Otherwise the colleague is asked to type the name.
5. **The pipeline itself was not run end to end in tests.** It is tested with
the pipeline, database and web lookups replaced by fakes. The first real run
should be the smoke test in §5, step 6.

View File

@@ -244,6 +244,59 @@ has always been. On, the same ladder runs with the app's vector and `text`
(no photo, so never server OCR), and the response is the identify shape
above. Not on the GET variant.
## When the product is not in the catalogue (capture-to-catalog)
With `ENABLE_CAPTURE_DISCOVERY=true` an unconfirmed identify no longer ends
at a best effort. The label is read into brand, product name, pack size and
category, and one of four things comes back in `discovery_status`:
| `discovery_status` | `matched_by` | What happened | What to show |
|---|---|---|---|
| `pending` | `discovery_pending` | Queued for the 11-stage pipeline. `results` is empty. | `provisional` + `discovery_message`; poll the job |
| `exists` | `label_exact` | The product is stored; the ladder just missed it. `results` holds it. | the product (confirmed) |
| `needs_input` | unchanged | The label could not be read, or the brand is not in the catalogue. | `discovery_message`: ask for the name, resend as `text` / `brand` |
| `busy` | unchanged | Rate limit (per device per hour) or full queue. | `provisional` + `discovery_message` |
`label_exact` is a confirmed answer. `discovery_status` is `null` when the
answer was already confirmed, or the feature is off.
```jsonc
{
"matched_by": "discovery_pending",
"results": [],
"discovery_status": "pending",
"discovery_job_id": "3f2a…", // 32 hex
"discovery_message": "Not in the catalog yet - adding it now. Details will fill in within a few minutes.",
"provisional": {
"brand": "Godrej", "product_name": "Godrej Fab Detergent Powder", "size": "1kg",
"category": "Detergents & Fabric Care", "hsn_code": "3402", "gst_percent": 18,
"hsn_gst_needs_review": false,
"visible_in_search": false, // brand not in ACTIVE_BRANDS: stored, but search hides it
"source": "label"
}
}
```
**Polling.** `GET /api/search/identify/jobs/{job_id}` returns `status`:
`queued`, `running`, then one of `done`, `rejected` (the validation gate refused
it; `detail` says why), `failed`, `interrupted` (the server restarted; capture
again). On `done`, `product` is the stored catalogue card with
`validation_status` `needs_review`, and `retail_presence` is the live check
of whether a retailer lists that exact pack. One job runs at a time, and each takes
roughly a minute or more (web images, retail lookup, LLM description).
**Only known brands.** A brand that is neither in the brand registry nor an
existing brand table is never created from a label, so an OCR misread cannot
mint a brand. A product line on its own ("Fab") is not treated as a brand.
**The photo.** It is kept under `CAPTURE_DIR`. It becomes the product image
only when the web search found no acceptable image *and*
`CAPTURE_PUBLIC_BASE_URL` is set; it is then served at
`/api/search/captures/{job_id}.jpg` and its vector is stored at once, so the
next photo of that pack matches on the image rung.
Admins can list recent capture jobs at `GET /api/admin/captures`.
## Errors
| Code | Cause | What to do |