Wiser Retail Strategies | Wiser Solutions

We gave Codex a catalog. Here's what it missed.

Written by Héloïse Tobin | 5 août 2026

Every so often a CFO looks at the pricing intelligence line item on the budget and wonders if the team could just build it themselves. Point Claude or Codex at the catalog, have it scrape the same retailers and skip the vendor fee entirely. On paper it sounds like a problem AI should be good at solving.

So we tested it on our own catalogs to find out.

Why a missed match doesn't look like a missed match 

The tricky thing about a missed match is that it doesn't announce itself. There's no error message and no red flag in a dashboard telling you a competitor's listing slipped through. The SKU is simply absent and every decision that depends on it gets made without that information considered. That's the real risk in swapping a purpose-built platform for a DIY AI scrape. Instead of an error being caught by a dedicated team, you'll find out weeks later, when a competitor's price move on a SKU you never tracked has already eaten into your margin. 

Test One: 5,000 SKUs, two retailers, 3.5 hours 

For the first test we ran Codex against a 5,000-SKU catalog spanning two retailers and gave it the same job our production system handles every day: build a competitive pricing feed. The job ran for three hours and 23 minutes before it finished. Accuracy came in at 48 to 49%, compared with Wiser's production accuracy on similar catalogs at 97%. Completeness, meaning the share of listings that should have matched and did, landed in the mid-20s% against Wiser's 90%. That's not a small gap: well over half of the competitive activity on that catalog never made it into the report. 

Test Two: a real customer's catalog, run in "goal mode" 

The second test gave the AI even less structure. We handed Codex a real customer's 5,361-product catalog and let it run in an autonomous "goal mode" with a single instruction to build a competitive pricing report. It took three hours and twenty-three minutes. When we spot-checked twenty or so of the matches it found, every single one was correct. So the AI is genuinely capable of confirming that two listings are the same product once it finds them. 

But finding them turned out to be the hard part. Codex surfaced 2,660 matches on Amazon and 39 on Google. Wiser's system finds roughly 4,000 on Amazon and 2,000 on Google for that same catalog, which means the AI run caught around 2% of what our system catches on Google alone. It also ran into problems a spreadsheet-and-API approach won't warn you about ahead of time, including getting IP blocked on one retailer with no proxy setup to route around it. The job burned through 28 million tokens, and 23 million of those were cached, so even running it again on the same catalog wouldn't get much cheaper. 

The cost math doesn't hold up 

That job cost $23.74 to run once at direct API pricing. Run daily over a year and that comes out to roughly $8,000, for one catalog of just over five thousand products, at less than a third the coverage of a system built for this. And that number only covers the API bill. It doesn't include the engineering time spent maintaining proxy rotation, retry logic and rate limits, all work Wiser has already solved at scale. The AI-DIY math only looks appealing if nobody adds up what the missing coverage and the maintenance behind it costs. 

Why the gap exists 

Part of why the gap is this wide comes down to what matching two listings requires. It isn't a single lookup: it's attribute extraction, standardization, categorization and a comparison across every attribute to decide whether two listings are close enough to call the same product. Wiser runs token-based and vector-based matching in parallel and routes the uncertain cases through an LLM for review, because running that logic in sequence simply wasn't accurate enough. A coding agent told to "build a pricing report" doesn't come with any of that infrastructure built in. It can be pointed at the problem, but it can't be pointed at the scale. 

The honest answer on Build vs. Buy 

If you're weighing whether to build this in-house, the honest answer is that AI can replicate a slice of what a pricing intelligence vendor does, and on a small sample that gets manually checked it can even look accurate while doing it. What it can't yet replicate is coverage at scale, resilience against the blocking every retailer throws up, or a cost curve that stays flat instead of climbing with the token bill as your catalog grows. The real question isn't whether AI can find some competitive prices. It's whether "some" is good enough to make a margin call on and whether the SKUs it misses happen to be the ones costing you the most.