Notes

What the model should do and what the code should do

Reading shelf price tags from photographs produced a simple rule: let the model understand the image, let the code verify the numbers.

28.07.2026 1 min read

Collecting competitors' prices means, in practice, photographing a shelf and then typing everything out. The obvious answer is to let a vision model read it. The first version did exactly that and worked surprisingly well, right up until we started checking the results.

The model sometimes read the price correctly but guessed the barcode. Not misread — guessed, and plausibly. A thirteen-digit number that looks like an EAN and is not one cannot be distinguished from a correct one in a CSV file.

The fix was not a better model but a division of labour. Barcodes are decoded by a specialised library straight from the image, deterministically and with an unambiguous outcome: either it reads the code or it says it cannot. The model takes on what it is irreplaceable at — working out which of the three prices on the tag is the valid one, whether it is a promotion, and which text is the product name.

Above that sits a third layer, validation. An EAN check digit can be verified by arithmetic. A price outside a sensible range is suspect. If a promotion is detected both by colour and by text and the two disagree, the record is flagged as uncertain rather than guessed.

The result is boring, which is the point. The deterministic layer is covered by a hundred and fifty tests that run offline with no paid calls. The model is called once per photograph and its output is always verified.

The generalisation we take from this: when something in a task can be verified by computation, never leave it to the model. Not because it cannot do it, but because the one time it fails, you will not notice.