Skip to content

We gave Codex a catalog. Here's what it missed.

Avatar for Héloïse Tobin

Sr. Director, Product Marketing | Wiser

Published

Duration

6 min read time

Last Updated: August 9, 2026

Every so often, a CFO looks at the pricing intelligence line item and asks a reasonable question: could we just do this with AI? 

Point an AI agent at the catalog, send it to Amazon and Google Shopping, and ask it to bring back the best matches. It sounds straightforward. The agent can browse, reason and output a spreadsheet. On paper, that looks close to a competitive price monitoring workflow. 

So we tested that assumption on a live catalog to find out what actually happens when a general-purpose AI workflow is asked to perform SKU-level price matching at production scale. 

The Test: 5,360 SKUs, Two Retailers, 3 Hours 23 Minutes 

We ran Codex against a 5,360-SKU catalog and asked it to find exact matches on Amazon.com and Google Shopping. 

The output looked useful at first glance: Codex found 2,260 exact matches on Amazon and 39 exact matches on Google Shopping, for 2,299 exact matches in total. It also flagged another 1,443 potential Amazon matches and 38 potential Google Shopping matches. 

Wiser's system found roughly 5,344 Amazon matches and 3,183 Google Shopping matches for the same catalog. 

On exact match rate, Codex covered about 42% of the catalog on Amazon and 0.7% on Google Shopping. Wiser's match-rate KPI for this same catalog was about 99% on Amazon.com and 59% on Google Shopping. 

Put another way: across the two retailers Codex returned 2,299 exact matched listings compared with Wiser's 8,527. That is roughly 27% of Wiser's matched-listing coverage. 

Why Missed Matches Are So Dangerous  

The tricky thing about a missed match is that it often does not look like a missed match. It looks like a product with no competitive activity. 

That is the dangerous part: if your competitor is undercutting you on a high-volume SKU and your workflow simply fails to find the match, your dashboard will not show a warning. 

That silence can lead teams to keep prices unchanged, miss margin pressure or assume they are alone in the market when they are not. In competitive price monitoring, missing data is not neutral. Quite the opposite in fact: it completely changes the decision the merchant makes. 

Accuracy Still Needs Independent QA  

After the match discovery run, we asked Codex to evaluate the accuracy of its own matches. That is useful as a rough signal, but it is not the same as an independent quality-control process because the same general workflow and AI model (GPT 5.6 Sol) are being used to judge the output it produced. 

Even with that limitation, the estimated accuracy came in at about 48%. 

At Wiser, we use AI in a more controlled way. Matching and validation are separate processes: one layer creates candidates, another evaluates match quality and uncertain cases can be routed through additional review. For this catalog specifically, Wiser's measured match accuracy was 98.4%. 

That is not a small gap. A generic AI workflow missed a large share of the available competitive activity and still needed significant review on the matches it did find. 

Why the Gap Exists 

The gap exists because product matching is not a single prompt, it's a production system with multiple AI and rules-based layers. 

Wiser is AI-heavy, but it is not just an LLM prompt pointed at a retailer website. Candidate generation runs through token matching and vector similarity in parallel; pre-validation rules confirm, reject or escalate candidates; and LLM validation is used for uncertain cases where language understanding adds value. 

Conceptual matching architecture

This is the important distinction: the model is swappable, but the matching system, evaluation stack, and guardrails remain controlled. 

That architecture lets Wiser use the latest state-of-the-art models where they help, without making the entire workflow dependent on one model's behavior. Model selection can be changed through configuration, but the surrounding guardrails - candidate rules, QA thresholds, evaluation checks, exception handling and monitoring - stay in place. 

The independent accuracy and completeness stack is the key distinction. It uses a different model and prompt from the matching flow, so match generation and match evaluation are not the same judgment wearing two hats. When a model changes, the eval stack helps catch drift before the change becomes production behavior. 

A coding agent can help assemble a one-off workflow, but it is not built with this kind of persistent product graph, retailer-specific matching logic, independent evaluation, exception queues, or match-rate monitoring. Those are the pieces that make the difference between an interesting spreadsheet and a pricing dataset a merchant can trust. 

Potential matches are a good example. They can be useful for triage, but they are not production-ready competitive prices. Someone still has to decide which candidates are true matches, which are false positives, and which should be suppressed or escalated. 

The Cost Math Doesn't Hold Up 

That single Codex run cost about $23.74 in model usage under the assumptions we used for the test. Running the same job once a day for a year would be about $8,665 before any engineering time, QA labor, proxy/browser infrastructure, retries, monitoring or maintenance. 

The token bill also already benefited heavily from caching. The session used about 28.52 million tokens in total, including about 26.95 million cached input tokens. A production process cannot evaluate cost from the successful run alone; it also has to account for failed searches, blocked pages, retailer layout changes, re-runs, review queues, and ongoing maintenance. 

That is the business problem. The DIY approach is not free. It is a partially automated workflow with material coverage gaps, accuracy risk, and operational overhead. 

Build vs. Buy Is the Wrong Frame

The question is not whether AI can help with price monitoring. It can, and it should. 

The better question is whether a general-purpose AI coding workflow can replace a dedicated comparative price monitoring system. For SKU-level competitive pricing, our test suggests the answer is no. 

A retailer does not need a spreadsheet that looks plausible. It needs trusted, repeatable market coverage; clear match-rate and accuracy metrics; and a workflow that keeps working after the first run. 

That is where dedicated tools matter. Wiser combines automation, AI-assisted matching, retailer-specific logic, human QA where it matters and production monitoring so the output can be used in real pricing decisions. 

For this catalog, Wiser delivered far broader coverage and much higher measured accuracy: about 99% Amazon match-rate coverage, 59% Google Shopping match-rate coverage, and 98.4% measured match accuracy. 

Wiser's production matching accuracy runs at 97%, with completeness in the 90s%, validated against scaling samples per catalog. If you want to see what full coverage looks like on your own catalog, talk to us!

FAQs

Not by itself. AI can assist with matching, review, and workflow automation, but reliable price monitoring also requires retailer-specific extraction, match-rate tracking, QA controls, and operational monitoring.

Codex's self-evaluation estimated match accuracy at about 48%. A small manual spot check did not find false positives, but that was not large enough to support a broad accuracy claim.

For this catalog, Wiser measured 98.4% match accuracy. Wiser's broader production benchmark runs at 97%+ accuracy, supported by ongoing validation and QA processes.

The single run cost about $23.74 in model usage under the test assumptions. Running it daily for a year would be about $8,665 before engineering time, review labor, infrastructure, retries, and maintenance.

Match rate measures how much of the catalog received a match. Accuracy measures whether the matches found are actually correct. A useful competitive pricing system needs both: high coverage and high confidence.

Wiser was built for this.

Blending AI with proven logic, Wiser turns billions of data points into fast decisions around pricing and execution.

See What's Possible
  • 10B+

    Products tracked

  • 4M+

    Prices recommended

  • 600K+

    Stores monitored

Trusted by brands that lead in every channel

  • Electrolux Logo
  • Nespresso Logo
  • Samsung Logo