This is the infrastructure every AI assistant story assumes already exists
Every grocery AI announcement in the past month, from conversational shopping assistants to ChatGPT plugins, quietly assumes the retailer's product catalog is clean, complete, and structured enough for an AI system to reason over. That assumption is usually wrong. Most grocery catalogs are a patchwork of vendor-supplied descriptions, inconsistent categorization, and missing attributes accumulated over years of manual entry. Hy-Vee's move to structure nearly 265,000 products across more than 1,200 attribute types is the unglamorous work that has to happen before any of the flashier AI features actually work well.
Averaging seven attributes per SKU across 5,800 active categories is a meaningful density. It is the difference between a search function that can only match on product name and one that can answer a query like find a low-sodium option similar to this brand, because sodium content, dietary flags, and comparable-product relationships are all now structured fields rather than buried in unstructured text. That distinction determines whether an AI shopping assistant feels genuinely useful or feels like a chatbot wrapped around a keyword search.
Choosing a data infrastructure vendor over a chatbot vendor was the right sequencing
Hy-Vee built this with Prodx through its Aisles app, a decision that prioritizes owning the data layer and digital merchandising controls directly rather than outsourcing decision logic to a third-party AI assistant platform. Prodx CEO Adam Christmann's comment that Hy-Vee's platform shows what the infrastructure can do at scale is a vendor talking his own book, but the underlying architectural choice is sound regardless of the source. Retailers that keep AI-relevant product data and merchandising control in-house retain more flexibility to swap or add AI-facing tools on top later.
Compare this to retailers that have gone straight to deploying a conversational AI assistant without first auditing catalog quality. Those deployments tend to produce impressive demo moments and disappointing real-world query success rates, because the assistant is only as good as the data it can retrieve. Hy-Vee's sequencing, data foundation first, AI features second, is the harder and slower path, but it is the one that scales without requiring a rebuild in 18 months when the assistant's limitations become obvious to customers.
Substitutions and predictions are where the ROI shows up first
The features Hy-Vee is layering on top, a predictions function that anticipates needs before a search happens, and a substitutions experience for out-of-stock items, are two of the highest-value applications of structured product data in grocery today. Substitution logic in particular has a direct and measurable revenue impact, since a poor substitute suggestion during an out-of-stock event is one of the most common reasons online grocery orders get abandoned mid-checkout or generate complaints and refund requests afterward. Instacart and other delivery platforms have spent years and considerable engineering effort trying to solve exactly this problem, with mixed results across categories.
Because Hy-Vee owns the underlying attribute data rather than depending on a third-party algorithm's black-box substitution logic, it can tune substitution rules directly against its own margin and inventory priorities rather than accepting whatever a delivery platform's generic model decides is a reasonable swap. That is a real competitive advantage over retailers who rely entirely on marketplace or third-party delivery-platform substitution logic they do not control and cannot adjust for their own category economics. Owning the data layer means owning the economics of the substitution decision, well beyond just shaping the customer experience of receiving it at the doorstep.
This Is a Scale Story for a Regional Grocer
Hy-Vee operates more than 560 stores across eight Midwestern states with annual sales exceeding 14 billion dollars, a large enough footprint that this data infrastructure investment functions as the backbone for a genuinely large-scale grocery operation, concentrated regionally rather than spread coast to coast. This is a serious commitment of budget and engineering time, not a boutique pilot confined to a handful of flagship stores while the rest of the chain runs on legacy catalog data. Regional grocers have historically operated with smaller technology budgets than the Krogers and Albertsons of the industry, which makes Hy-Vee's willingness to fund a full catalog restructuring across its entire footprint particularly notable.
It also suggests the cost of this kind of foundational data work has come down enough that mid-size regional retailers can now credibly undertake it, rather than it staying reserved for chains with the largest IT budgets and the deepest bench of in-house data engineers. CIOs at similarly sized regional retailers should treat this as concrete evidence that catalog-level data infrastructure investment is achievable at their scale and their budget, and not only at the scale of the top five national grocery chains that usually headline this kind of coverage.
The lesson for every retailer chasing an AI shopping assistant headline
The retail AI news cycle rewards the visible layer: chatbots, voice assistants, conversational carts. Hy-Vee's announcement got comparatively little attention because structured product attributes do not photograph well and do not generate a demo video. But every retailer currently evaluating an AI shopping assistant vendor should be asking a more basic question first: does our product catalog have the attribute density this assistant needs to actually answer the questions customers will ask it.
If the honest answer is no, the higher-leverage investment this year is closer to what Hy-Vee just did than to licensing another conversational AI layer. The retailers who get this sequencing right will end up with AI features that compound in usefulness as the catalog improves further. The ones who skip straight to the chatbot will spend the next two years discovering, expensively and publicly, why their assistant keeps giving customers wrong or incomplete answers.

