AI search optimisation Shopify

llms.txt: Emerging Standard Or Comfort Blanket? An Evidence Check

llms.txt: Emerging Standard Or Comfort Blanket? An Evidence Check

I build a GEO tool. I have a commercial interest in you believing that AI search optimisation is real, worth doing, and worth paying for. So I want to be scrupulous about this one, because llms.txt is the closest thing the field has to a shibboleth — implement it and you’re taking AI search seriously, don’t and you’re behind — and I don’t think the evidence supports that framing.

Here’s the honest version.

What llms.txt proposes

The idea is straightforward and, on its face, sensible. A plain-text markdown file at your domain root — `/llms.txt` — containing a curated, machine-friendly summary of your site: what it is, what the key pages are, where the important documentation lives, expressed cleanly without navigation, scripts or markup.

The reasoning by analogy is to `robots.txt` and `sitemap.xml`. Those work. Both are conventions a site publishes at a known path that crawlers agreed to read. llms.txt asks for the same social contract from language models.

There’s a real problem underneath it. Web pages are a hostile format for extraction: navigation, cookie banners, related-product carousels, JavaScript-rendered content, boilerplate. A clean markdown summary genuinely would be easier to parse. The motivation isn’t silly.

The question that decides everything

Here’s the thing about robots.txt and sitemap.xml, though. They work not because they’re good ideas, but because **the major crawlers publicly committed to reading them**. Google documents its robots.txt handling in detail. Sitemap support is announced, specified and maintained.

So the only question that matters for llms.txt is: 

**has any major AI provider publicly confirmed they consume it?**

That’s the question to answer before you spend an afternoon on this, and it’s the question almost every post advocating llms.txt carefully doesn’t address. They discuss adoption — how many sites have published one — which is a measure of how persuasive the idea is, not of whether it does anything.

I’m not going to state a current answer here, because the position could change between my writing this and your reading it, and a confident claim that ages badly is exactly the disease I’m complaining about. What I’d urge instead is a five-minute check you can run yourself:

– Look at the published documentation for OpenAI’s, Google’s, Anthropic’s and Perplexity’s crawlers. Does any of them mention llms.txt as an input? Their crawler docs are public and specific about what they do read.
– Check your own server logs. If AI crawlers are fetching `/llms.txt`, you’ll see the requests. This is the most direct evidence available and almost nobody bothers.

That second one is genuinely the best test in existence right now, and it costs you a log query. If you have a file published and it’s never been requested, you have your answer for your site.

The evidence for and against a citation lift

Beyond documentation, the standard you’d want is a controlled test: comparable sites or comparable sections, llms.txt on some and not others, citation share measured over a meaningful window.

I haven’t seen one done credibly. What circulates instead is case studies — “we added llms.txt and our AI citations rose 40%” — which are almost always confounded. Sites that publish llms.txt are typically sites actively working on AI visibility, which means they’re simultaneously improving structured data, publishing more, tightening content and building mentions. Attributing the outcome to the text file is the weakest available explanation.

Contrast that with what *is* documented. Shopify Catalog syndicates product data to ChatGPT, Copilot, Google AI Mode and the Gemini app, with real-time price and inventory verification — that’s a confirmed, specified pipeline into named AI surfaces. Product schema is documented, versioned and machine-readable, with published requirements. Shopify’s own semantic search is documented as drawing on product descriptions and image data.

Those are mechanisms with documentation behind them. llms.txt, currently, is a proposal with adoption behind it. Those are very different kinds of evidence and they shouldn’t be weighed equally just because both feel like “AI SEO.”

What it costs you to add anyway

Having been sceptical, let me steelman it, because the cost side genuinely matters.

Publishing an llms.txt is close to free. It’s a static markdown file. There’s no performance cost, no crawl-budget cost, no risk of a penalty, no ongoing maintenance beyond keeping it roughly current. If a provider does adopt it, you’re already there.

There’s also a modest side benefit that has nothing to do with AI: writing one forces you to articulate what your site is and which twenty pages actually matter. On more than one occasion that exercise has been more useful to a client than the file.

So: a genuinely cheap option on an uncertain future. I have no problem with anyone publishing one.

My problem is with what it displaces. An afternoon is an afternoon, and if the choice is between llms.txt and populating the fabric-weight metafield across 200 products, the second one has a documented mechanism, a documented consumer, and a measurable effect on your own site search and filters before you even get to AI.

What to do instead if you only have one afternoon

Ranked by evidence quality, strongest first:

1. **Populate structured product attributes** — metafields, not prose. Documented input to Shopify’s own search, Catalog syndication, and every filter on your site. Blank fields exclude you from filtered queries entirely; there’s no partial credit.
2. **Validate Product and Offer schema** on your top templates, confirming price and availability match the live page. Documented, specified, verifiable in minutes.
3. **Confirm your Catalog eligibility and syndication status.** A documented pipeline to four named AI surfaces. Worth knowing whether you’re actually in it.
4. **Rewrite your top product descriptions as factual declarative sentences**, with the brand voice kept but confined. Extractable content is the entire mechanism.
5. **Fix your top zero-result site search queries.** Immediate revenue effect, and it improves the same vocabulary layer everything else reads.
6. **Then, if you like, publish an llms.txt.** It’s cheap, it might pay off, and by this point everything it would point at is actually worth pointing at.

Verdict:

llms.txt is a reasonable proposal to a real problem, published by people acting in good faith, with a plausible mechanism and — as far as I can establish — no confirmed consumer among the major providers and no controlled evidence of effect.

That makes it a cheap bet, not a best practice. Publish one if you want the option. Check your logs to see if anything ever asks for it. Don’t let anyone tell you it’s table stakes, and be wary of anyone who presents it as the thing standing between you and AI visibility — because that framing usually comes from someone who’d rather sell you a file than an audit of your product data.

The unglamorous stuff still wins. It has for twenty years of search, and nothing in the last two has changed it.

If there’s a general lesson here beyond one text file, it’s about how to evaluate every new AI SEO tactic that will be pitched to you over the next eighteen months — and there will be many. Ask three questions. Is there a documented consumer, named by the provider rather than inferred? Is there evidence of effect that isn’t a confounded case study? And what does doing it displace? Most of what’s currently being sold as GEO fails at least two of those. The tactics that pass are, disappointingly, the same fundamentals that have always worked: accurate data, specific writing, and being the sort of business other people mention.

*Sources: Shopify, “Agentic Commerce on Shopify: How It Works” (2026); Shopify Help Centre, Search & Discovery documentation. Before publishing, verify current llms.txt positions directly from OpenAI, Google, Anthropic and Perplexity crawler documentation, and cite any adoption figure only with a named crawl and date.*

Requires fresh primary research at draft time: (a) confirmed statements from OpenAI, Anthropic, Google and Perplexity on whether llms.txt is consumed; (b) current adoption counts from published crawls; (c) any controlled test showing citation lift. Contrast with what IS documented: Shopify’s Catalog syndication and Product schema are confirmed machine-readable inputs to AI shopping surfaces (Shopify, 2026). Do not publish an adoption percentage without a named crawl and date.

✓ No theme edits

✓ Shopify billing

✓ First audit < 5 min

✓ Cancel anytime

Discover more from GEOptimisation

Subscribe now to keep reading and get access to the full archive.

Continue reading