What the instrument doesn't look at
If the product is going to say every day where things are cheaper, it had better be right. I spent this week measuring how often it isn't.
Having the product speak every day comes with an uncomfortable consequence: now it has to be right every day. I built my own weekly list, compared it, and the savings came out small. Instead of taking that at face value I asked the question I didn’t want to ask: what if the savings are small because the system isn’t seeing every alternative that exists? If a product’s card is missing precisely the store where it’s cheapest, the comparator isn’t lying — it simply never found out. That’s the worst kind of error, because it leaves no trace: nobody complains about a deal they never saw.
Answering that needed an instrument, and that’s where the lesson I’m least proud of and got the most out of showed up. There are two opposite errors — cards missing a product that should be there, and cards that merged things that aren’t the same — and it is dangerously easy to improve one number by quietly making the other worse. My first design for the measurement only looked where I believed the problem lived. I caught it in time: the sample was redrawn across the entire catalog, close to 57,000 cards, and every card drawn gets asked both questions. The result justified the change, because the worst segment of all turned out to be exactly the one my first design would have declared healthy.
The draw came to 1,048 cards from across the catalog, reviewed one by one, with a three-point margin at 95% confidence. Out of every hundred cards, twenty-one are incomplete and one has something inside it that doesn’t belong. In other words: the catalog isn’t badly merged, it’s badly completed. And the detail that gave me the most to think about concerns the previous instrument, the one I had been trusting: it was seeing fourteen percent of what was actually there. An instrument that can only find one class of error ends up certifying the class it doesn’t look at.
Which leaves the rule I don’t intend to move: a number that falls short of its target gets raised by fixing the product, never by loosening the yardstick that measures it. What comes next is closing that gap — including recovering a piece of data that one of the chains doesn’t publish and the others do, and which today keeps thousands of cards from finding their twin. Every point recovered there is one more store the comparator can see, and money that shows up on somebody’s screen.