Image vector embedding
This commit is contained in:
@@ -477,13 +477,15 @@ def _brand_defaults(brand_parent: str) -> Dict[str, Any]:
|
||||
"""
|
||||
# Two reads on purpose, and the split matters on a memory-capped host.
|
||||
#
|
||||
# `_brand_sample` is SELECT * limit 1 - one row, embedding and all, because
|
||||
# the fields it seeds (category, price band, size) need the whole row.
|
||||
# `_brand_sample` is one full row (limit 1), because the fields it seeds
|
||||
# (category, price band, size) need the whole row. The product readers now
|
||||
# project the vector columns out, so "full" no longer means "with the
|
||||
# embedding".
|
||||
#
|
||||
# The consensus read is 300 rows, so it takes only the two columns it
|
||||
# actually inspects. Measured: SELECT * over 244 rows costs 3.0 MB, of
|
||||
# which 4.7 KB per row is an embedding string nothing here reads. Two named
|
||||
# columns is roughly 50 KB for the same rows.
|
||||
# actually inspects. Measured when reads were still SELECT *: 244 rows
|
||||
# cost 3.0 MB, of which 4.7 KB per row was an embedding string nothing
|
||||
# here reads. Two named columns is roughly 50 KB for the same rows.
|
||||
rows = consensus_rows(brand_parent, ["fssai_license", "providers"],
|
||||
limit=_CONSENSUS_SAMPLE_LIMIT)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user