Home Bilby — store monitor Free store scan Free tools AI apps Navaal SEO app Blog Services Studio Contact Sign in

// Research · Product data

We checked the product data on 409 live Shopify stores

By Published Updated 7 min read

Last updated 24 September 2026 — the audit month now sits beside every statistic. Public web only — no logins, no admin access, no store data. Every number below is reproducible from the method at the end.

71.9% of 409 live Shopify storefronts had at least one actionable product-data problem — but that headline is carried by two content-quality checks, and it is the wrong number to plan around. Strip out "a product is missing its product_type" and "a description is under 120 characters" and the figure falls to 36.2%. Narrow it again to the attributes that AI shopping feeds actually require (as of September 2026) — price, image, policy pages — and it is 22.7%. The median store has one finding, not a pile of them. And the thing everyone writes about — AI crawlers being blocked in robots.txt — turned up on 2 stores out of 409.

The four numbers, with the caveat attached

We audited 409 live Shopify storefronts from the public web in September 2026. A "finding" means a check failed in a way a merchant could act on. Here is the headline next to the sensitivity analysis, because publishing the first row on its own would be misleading:

409 live Shopify storefronts, public-web audit, September 2026
How you define "actionable finding"Share of stores
All nine checks71.9% (294/409)
Excluding missing product_type and thin descriptions36.2% (148/409)
Only attributes AI shopping feeds require, plus policy pages22.7% (93/409)
Only structured data (missing or duplicated)15.9% (65/409)
Only crawler access0.5% (2/409)

Median findings per store in the September 2026 sample: 1 (mean 1.28, maximum 6). Stores with zero findings: 28.1%. And 35.7% of all stores had no finding except a missing product_type and/or a short description.

Across 409 live Shopify storefronts audited from the public web in September 2026, 71.9% had at least one actionable product-data finding, but only 22.7% were missing an attribute that AI shopping feeds actually require, and the median store had one finding.

What fails, ranked

Each check is scored only on the stores where it could be measured in the September 2026 audit, so the denominators differ.

Nine checks across 409 stores, September 2026
CheckFailedOfRate
Missing product_type on at least one sampled product17338544.9%
Description under 120 characters on at least one product16938543.9%
No price on at least one product4038510.4%
No image on at least one product373859.6%
A policy page missing (terms, privacy or refund) — checked September 2026344098.3%
Duplicate Product JSON-LD (two or more top-level blocks)334028.2%
No structured product markup of any kind324028.0%
Canonical missing, malformed, or pointing off the product34020.7%
Crawler blocked by robots.txt24090.5%

At the individual product level, across the 3,784 products sampled in September 2026:23.9% had no product_type, 17.5% had a description under 120 characters, 3.0% had no image and 3.8% had no price.

A note on the top row, because it is the one most likely to be oversold: product_type is Shopify's own taxonomy field. It is not on the list of attributes OpenAI's product feed specification requires. Treating it as a blocker would overstate what it does. It is worth filling in; it is not the reason a product does or does not show up.

The two samples agree, which is the most useful part

We built the sample two different ways on purpose, from unrelated sources, to see whether the result held.

Sample A — larger storesSample B — the long tail
How it was builtShopify storefronts inside the Tranco top 49,560Random sample from 3,392 verified live Shopify stores in a public list
Stores audited (September 2026)186223
At least one finding72.0%71.7%
Median findings11
Zero findings28.0%28.3%

Sample A skews to large, well-resourced merchants and sample B to small ones. We expected the bigger stores to look visibly cleaner. They did not — the two land within 0.3 percentage points of each other.

The result we did not expect: crawler blocking is basically not a thing

A lot of advice right now is about stopping AI crawlers from being blocked. We checked robots.txt on all 409 stores, agent by agent, for OAI-SearchBot, PerplexityBot, Claude-SearchBot, bingbot and Googlebot.

  • Two stores out of 409 disallowed any of them.
  • Not one store blocked an AI crawler while allowing Google. One blocked everything including Googlebot; one blocked bingbot only.
  • As of September 2026, only about 1.6% of robots.txt files mention an AI crawler group at all. Almost nobody has edited the default.
  • A further nine stores served a blanket Disallow: / — and every one of them also had no reachable product page. Those are closed or password-protected storefronts, not misconfigured open ones, so they are excluded from the base.

In a September 2026 audit of 409 live Shopify storefronts, only two disallowed any major search or AI crawler in robots.txt, and none blocked an AI crawler while still allowing Googlebot.

We also checked the agent-facing endpoints as a control in September 2026:/llms.txt returned 200 on 100% of stores, /agents.md on 99.5%, and /.well-known/ucp on 100%. Shopify serves all three on every live store by default. If someone offers to create those files for you, they are offering to do something the platform already did.

One thing we could not measure at all

Barcodes. A product's barcode is not exposed in a Shopify store's public product JSON — it is admin-only. So the barcode completeness rate across the market is unknown to us, and it is unknown to anyone else auditing from the outside. We would rather say that than estimate it.

Method, and three mistakes we made first

For each store, paced at roughly two requests per second and identified in the user agent: robots.txt; the public product JSON, sampling up to ten products with a seed derived from the domain so the sample is reproducible; one product page; five crawler user agents against the homepage and a product URL; three policy pages; three agent endpoints. Stores with no reachable product page and no product JSON were excluded as unusable (37 of 446) rather than counted as passes.

Three things went wrong in the first pass. Each would have produced a publishable but false number, and — worth noting — every one of them made the market look more broken than it is:

  1. DNS starvation faked a clean crawler result. A parallel scan saturated the resolver and 137 of 250 robots.txt fetches returned nothing. The analysis treated "could not read" as "no rule, therefore allowed", and reported a 0% block rate over a half-empty denominator. Re-fetched on its own with retries, 248 of 250 returned 200.
  2. "No product schema" was wrong by a factor of seventeen. In the September 2026 first pass, 68 stores had no Product JSON-LD, which read as 30.6%. On a re-check, 58 of those 68 emit itemtype="schema.org/Product" as microdata, which Google accepts. The true rate is 8.0%.
  3. Canonicals and duplicate schema were both over-counted. A canonical pointing at /en-us/products/x is correct for a locale storefront, not broken, and a Product plus a ProductGroup is Google's recommended variant markup, not a duplicate. Re-classified, the canonical failure rate fell from 7.7% to 0.7%.

Two caveats we are leaving in rather than burying. HTTP status codes returned to a crawler user agent came from our own IP, which is not a verified crawler IP, so they are not evidence of what a real crawler would see — the robots.txt directives are the finding, and no number above depends on those status codes. And static fetches do not execute JavaScript, so a store that injects its structured data client-side would read as missing; the microdata re-check in point 2 is what bounds that error.

What we would actually do with this

If you run a store, the useful reading of the September 2026 numbers is not "71.9% of stores are broken". It is that the median store has one thing wrong, that the most common problems are content-shaped rather than technical, and that the technical failure everyone is currently being sold — blocked AI crawlers — is real but rare enough that it should be a thirty-second check, not a strategy.

We build Navaal SEO, a Shopify app that audits a catalogue and writes the missing product content in your own brand voice, with nothing published until you approve it. It is the reason we ran this study: we wanted to know whether the audit would have anything to say to a typical store before we built more of it. For roughly seven stores in ten it does — and for the rest, the honest answer is that their product data is fine.