Image vector embedding

This commit is contained in:
sriram
2026-09-16 16:33:39 +05:30
parent ce4fa70dee
commit 9c8dbf1759
10 changed files with 1493 additions and 17 deletions

View File

@@ -477,13 +477,15 @@ def _brand_defaults(brand_parent: str) -> Dict[str, Any]:
"""
# Two reads on purpose, and the split matters on a memory-capped host.
#
# `_brand_sample` is SELECT * limit 1 - one row, embedding and all, because
# the fields it seeds (category, price band, size) need the whole row.
# `_brand_sample` is one full row (limit 1), because the fields it seeds
# (category, price band, size) need the whole row. The product readers now
# project the vector columns out, so "full" no longer means "with the
# embedding".
#
# The consensus read is 300 rows, so it takes only the two columns it
# actually inspects. Measured: SELECT * over 244 rows costs 3.0 MB, of
# which 4.7 KB per row is an embedding string nothing here reads. Two named
# columns is roughly 50 KB for the same rows.
# actually inspects. Measured when reads were still SELECT *: 244 rows
# cost 3.0 MB, of which 4.7 KB per row was an embedding string nothing
# here reads. Two named columns is roughly 50 KB for the same rows.
rows = consensus_rows(brand_parent, ["fssai_license", "providers"],
limit=_CONSENSUS_SAMPLE_LIMIT)