← All articles

1 September 2026

llms.txt File for AI Crawlers: A UK Website Owner's Guide

An llms.txt file for AI crawlers is a plain markdown file placed at your domain root (for example, yourbusiness.co.uk/llms.txt) that summarises your site's key pages and links for large language models to read. On its own it doesn't force any AI system to fetch or cite your content, but paired with robots.txt rules for GPTBot, PerplexityBot and Google-Extended, it gives UK website owners the clearest available way to signal what they want AI answer engines to access, ignore, or use for training and citation.

What Is an llms.txt File, Exactly?

An llms.txt file is a lightweight markdown document, not a technical access-control mechanism. It typically contains:

  • A short description of the business, written in plain language
  • A curated list of your most important pages, with one-line descriptions
  • Links to documentation, FAQs, or product pages you specifically want an AI system to understand

It sits at the root of your domain, the same location as robots.txt, so it's easy for any system that supports the convention to find. The important caveat for UK site owners: llms.txt is an emerging, voluntary convention. Not every AI crawler currently fetches or honours it the way established systems honour robots.txt, which has been a widely respected web standard for decades. Treat llms.txt as a helpful signal and a tidy reference document, not as a guaranteed gatekeeper.

Which AI Crawlers Actually Matter for a UK Site?

Most of the practical control happens through your existing robots.txt file, using specific user-agent tokens for each AI crawler. Here's what UK site owners are most likely to encounter:

Crawler Operator What it's for robots.txt token
GPTBot OpenAI Fetches pages that may inform ChatGPT's responses User-agent: GPTBot
PerplexityBot Perplexity Fetches pages used in Perplexity's live answers User-agent: PerplexityBot
Google-Extended Google Controls use of your content for Gemini and AI features, separate from classic Googlebot indexing User-agent: Google-Extended
ClaudeBot / anthropic-ai Anthropic Fetches pages that may inform Claude's responses User-agent: ClaudeBot

A crucial distinction for anyone running a UK business site: Google-Extended does not control whether your pages appear in ordinary Google search results. That's still governed by Googlebot. Google-Extended specifically governs whether your content can be used for Google's generative AI features, so you can, in principle, stay fully indexed in classic search while opting your content out of that particular AI use case, or vice versa.

How Do You Actually Set These Rules on a UK Website?

You control AI crawler access through the same robots.txt file most UK sites already have, whether it's hosted on WordPress, Shopify, Webflow, or elsewhere. A basic pattern looks like this:

  1. Open or create robots.txt at yourdomain.co.uk/robots.txt
  2. Add a block per crawler, for example:
  • User-agent: GPTBot
  • Disallow: /
  1. Repeat for PerplexityBot, Google-Extended, or ClaudeBot, using Allow: / instead of Disallow: / if you want that crawler to fetch your pages
  2. Save, publish, and re-check the live file at your domain to confirm it's being served correctly
  3. Optionally, add a separate llms.txt file at the root summarising the pages you most want an AI system to understand, distinct from the pages you'd rather it ignored

Most UK content management systems and platforms let you edit robots.txt directly or through an SEO plugin. If you're on a hosted platform with limited file access, check your platform's specific settings panel for a robots.txt or crawler-control option before assuming you can't edit it.

Should a UK Business Block or Allow AI Crawlers?

This depends entirely on your goal, and there's no universally correct answer:

  • If you want to appear in AI-generated answers (ChatGPT responses, Perplexity results, Google AI Overviews), you generally need to allow the relevant crawlers, since a system that can't fetch your page usually can't cite it as a source.
  • If you're concerned about your original content being used to train models without appearing as a live citation, Google-Extended and equivalent tokens let you separate "can this inform an AI answer right now" from "can this be used in longer-term model training," though the practical distinction varies by operator and isn't identical for every crawler.
  • If you run a UK business that depends on discoverability, such as a Bristol-based service provider, a Shopify store shipping across England, Scotland and Wales, or a Manchester agency managing several client sites, blocking AI crawlers outright works against the current shift where customers increasingly ask AI systems directly for recommendations instead of clicking through classic search results.

There is no setting that guarantees a specific AI engine will cite you, and no crawler configuration can promise immunity from Google's ranking or indexing decisions. What crawler settings and llms.txt give you is control over access, not a guarantee of outcome.

How Does Crawler Access Relate to Getting Actually Cited?

Allowing GPTBot, PerplexityBot, and Google-Extended to fetch your pages is a precondition, not a strategy. Being fetchable doesn't make your content citable. Answer engines extract and quote sources that answer a specific question clearly, in a self-contained way, without padding. This is where content structure matters as much as crawler permissions: FAQ-style sections, direct answers stated plainly near the top of a page, and specific, verifiable detail are what actually earn the citation once a crawler can reach the page. This is the core discipline behind answer-first content writing, where every piece is built to be lifted and quoted, rather than written for keyword density aimed at a results page.

What Does This Mean for Agencies and Multi-Site UK Operators?

Agencies managing several client domains face a practical problem: crawler settings and llms.txt files have to be checked and maintained per site, and content strategy has to vary by client without angles or access bleeding between accounts. That's a configuration and content-production workload on top of everything else on a retainer, which is part of why Rankmoss's services run campaign and team management with per-project access control alongside the answer-first writing itself, so an agency can keep each client's crawler-friendly, AI-citable content separate and on schedule without manually rebuilding the process for every account.

FAQ

Does having an llms.txt file guarantee ChatGPT or Perplexity will cite my UK website?

No. An llms.txt file is a voluntary, still-emerging convention, and no AI answer engine guarantees citation based on its presence. It can help summarise your site clearly for systems that choose to read it, but the actual decision to cite a page depends on whether the crawler can access it and whether the content itself directly and clearly answers the question being asked.

Will blocking GPTBot or PerplexityBot hurt my normal Google search rankings?

No, blocking AI-specific crawlers like GPTBot, PerplexityBot, or Google-Extended does not affect classic Googlebot indexing or your standard search rankings, since these are separate crawlers with separate purposes. It will, however, mean those AI systems generally cannot fetch or cite your pages in their generated answers.

Do I need separate llms.txt files for different languages if my UK business trades internationally?

There's no fixed requirement, but if you serve multiple markets or languages, it's reasonable to reflect that structure in your llms.txt summary and in the underlying content itself. Businesses expanding into new markets often need multi-market, multi-language content generation so the pages an AI crawler finds are actually written for each audience, not just translated crawler permissions layered onto English-only content.

Is editing robots.txt for AI crawlers something a small UK site owner can do without a developer?

On most common platforms, yes. WordPress, Shopify, Webflow, and similar systems allow direct or plugin-based editing of robots.txt, and adding a few User-agent and Disallow or Allow lines for GPTBot, PerplexityBot, or Google-Extended doesn't require custom development. If your platform restricts file-level access, check its built-in SEO or crawler settings panel first.

How do I know if allowing AI crawlers is actually leading to citations rather than just extra crawler traffic?

The most reliable approach is tracking which of your pages and questions are gaining visibility over time using tools like Google Search Console and GA4, then directing future content toward what's working. This kind of analytics feedback loop is standard across Rankmoss's plans, which is worth reviewing if you want crawler access paired with content built specifically to be extracted once it's reachable.

This article was written and published by RankMoss.

See how it works →