LLMs.txt

LLMs.txt is a proposed Markdown file that gives AI systems a clearer guide to a website’s most important, useful, and authoritative content.

What is LLMs.txt?

Quick definition: LLMs.txt is a proposed web convention that lets site owners place a Markdown file at the root of a domain to help large language models, AI agents, and retrieval systems understand which pages matter most.

The file usually lives at a URL like this:

https://example.com/llms.txt

It may include a short site description, links to important pages, notes about content categories, and short explanations of why each linked resource is useful. The goal is to give AI systems a clean, curated map instead of making them infer a site’s purpose from navigation, templates, sidebars, archives, scripts, and footer clutter.

For writers, publishers, SEO teams, affiliate site owners, and content marketers, the file sits at the intersection of AI visibility, technical SEO, content strategy, and generative search. It may help AI systems understand what the site is about. It may also do very little if the systems you care about ignore it. Welcome to modern visibility work, where the map may or may not be read by the robot.

What it is used for

An AI-readable guide is used to point large language models and AI systems toward a site’s most useful content.

It can highlight:

  • Core pages
  • Glossary entries
  • Documentation
  • Product reviews
  • Buying guides
  • Comparison pages
  • Reference articles
  • Publisher information
  • Editorial standards
  • Affiliate disclosure pages
  • Policy pages
  • Important topic hubs

For Scribbright, the file could point AI systems toward pages about writing tools, editing software, desk gear, productivity systems, AI writing tools, affiliate reviews, and glossary definitions for working writers.

That matters because a site owner should not make a machine guess which pages best explain the site. Human readers should not have to guess either, but machines have even less patience and worse taste in navigation.

Why it matters

AI systems increasingly retrieve, summarize, cite, and recommend web content.

Traditional websites are built for human visitors and search crawlers. They include menus, images, scripts, ads, templates, category links, related posts, footers, tracking code, and other useful but noisy elements. Those pieces help the normal site experience, but they can make it harder for AI systems to extract the main editorial value cleanly.

A curated file can help communicate:

  • What the site covers
  • Who the site serves
  • Which pages are most important
  • Which resources are authoritative
  • How the content is grouped
  • Which pages may be useful for AI retrieval

This does not guarantee AI citations, traffic, inclusion, ranking, or recommendations. It is a signal. Signals can be useful. Guarantees are usually sold by people who also own suspiciously confident webinars.

Where the file goes

The file should usually live at the root of the domain.

For example:

https://scribbright.com/llms.txt

That placement mirrors files like robots.txt, which also live at the root. The file should be publicly accessible and written in clean Markdown rather than HTML.

A useful version should be concise enough for an AI system to scan, but specific enough to explain what matters. A bloated file that lists every page defeats the point. That is not curation. That is a sitemap wearing Markdown.

What goes into the file

A site guide should give AI systems a short, useful overview of the domain and its strongest resources.

Common elements include:

  • Site name
  • Short description
  • Primary audience
  • Main content categories
  • Important URLs
  • Short page descriptions
  • Reference pages
  • Review hubs
  • Buying guide hubs
  • Comparison pages
  • Editorial or disclosure pages
  • Optional usage notes

The descriptions should be specific. “This page contains information” is technically accurate and spiritually asleep.

A simple template

A basic Markdown guide might look like this:

# Site Name
Brief description of the site, its audience, and its main purpose.
## Core Pages
- [Homepage](https://example.com/): Short description of the site and its main purpose.
- [Reviews](https://example.com/reviews/): Reviews of tools, products, or resources for the site’s audience.
- [Glossary](https://example.com/glossary/): Definitions of important terms covered by the site.
## Key Resources
- [Important Guide](https://example.com/guide/): Description of what this guide explains and who it helps.
- [Comparison Page](https://example.com/comparison/): Description of what this page compares and why it matters.
## Policies
- [Affiliate Disclosure](https://example.com/disclosure/): Explanation of affiliate relationships and editorial policy.

The template should be adapted to the site. A software company, review site, glossary, publisher, documentation hub, and ecommerce site should not all use the same structure.

The point is to make the important content easier to find and understand, not to create another ritual file no one can explain.

LLMs.txt vs. robots.txt

LLMs.txt and robots.txt do different jobs.

A robots.txt file tells crawlers which parts of a site they are allowed or not allowed to crawl. It can include directives such as:

User-agent: *
Disallow: /private/

An AI-readable Markdown guide does not enforce crawl rules. It does not block bots. It does not grant or deny permission. It points systems toward selected content and explains why those pages matter.

A simple distinction:

  • robots.txt: crawl access guidance
  • Markdown site guide: curated content context for AI systems

One is a gate sign. The other is a map. Confusing them can create the sort of technical misunderstanding that later appears in a meeting with too many people.

How it differs from sitemap.xml

A sitemap.xml file lists URLs for search engines. It is usually more comprehensive and technical.

A curated AI guide should be more selective. It should highlight the pages that best represent the site’s authority, usefulness, editorial focus, product categories, or reference value.

A sitemap says, “Here are the URLs.”

A curated AI guide says, “Here is what matters and why.”

Both can help discovery, but they serve different readers. The sitemap is for crawlers. The Markdown guide is for systems that need context without wandering through every hallway.

How it differs from schema

Schema is structured data added to webpages to help search engines understand page types, entities, products, reviews, FAQs, breadcrumbs, articles, and other structured information.

A root-level Markdown guide is separate. It gives AI systems a curated overview of important site content.

Schema usually lives on or near individual pages. The site guide lives at the domain root. Schema helps describe a page. The guide helps select and explain the pages a site owner considers most useful.

For Scribbright, both can matter. A glossary entry may use FAQPage schema. The site guide may link to the glossary as a core reference hub. One marks up the page. The other curates the map.

How it differs from noindex

A noindex directive tells search engines not to include a page in search results.

The Markdown guide does not tell search engines or AI systems to exclude a page. It is not an indexing control. If a site owner wants a page excluded from search, that belongs in meta robots tags, X-Robots-Tag headers, or related indexing controls.

The guide is for directing attention toward useful pages. Noindex is for keeping pages out of search results.

Different jobs. Same internet. Easy to confuse if everyone is moving too quickly and saying “AI” every third sentence.

How it differs from an AI training opt-out

The file is not an AI training opt-out mechanism.

It does not, by itself, tell AI companies they may not train on a site’s content. It does not create a technical block. It does not guarantee that crawlers, models, agents, or search systems will obey anything inside it.

For crawl control, site owners may use robots.txt rules, bot-specific directives, server controls, access restrictions, legal terms, or other mechanisms depending on the goal.

The site guide is better understood as an inference-time aid: “Here is the content that best explains this site.”

That is useful. It is not the same as “stay out.”

How it differs from LLMs-full.txt

LLMs-full.txt is sometimes discussed as a companion file.

The general distinction is:

  • llms.txt: a concise index or guide to important content
  • llms-full.txt: a longer file that may include fuller page content for easier ingestion

The shorter file is the lightweight map. The longer file is closer to a bundled source document.

Not every site needs the fuller version. It may be useful for documentation-heavy sites, developer tools, product docs, structured reference libraries, or sites that want to provide more complete source material in one place.

For many publishers, a concise guide is enough to start. As usual, the answer depends on the site, the content, and whether anyone is maintaining the thing after the first burst of enthusiasm.

GEO and AI visibility

Generative engine optimization, or GEO, focuses on how brands, sites, products, and sources appear in AI-generated answers.

A curated site guide can support GEO by helping AI systems understand:

  • Site purpose
  • Core topics
  • Important pages
  • Content categories
  • Glossary hubs
  • Product review areas
  • Buying guides
  • Comparison resources
  • Publisher and editorial context

It will not make thin content authoritative. It will not fix weak positioning. It will not turn a vague review into a useful recommendation.

But it can make a strong site easier for AI systems to interpret. That is not magic. It is a table of contents for machines.

AEO and direct answers

Answer engine optimization, or AEO, focuses on making content useful for direct answers.

A root-level guide may support AEO by pointing AI systems toward pages that answer important questions clearly.

Useful candidates include:

  • Glossary pages with direct definitions
  • FAQ pages
  • Comparison pages
  • Buying guides
  • Product reviews
  • Reference articles
  • Documentation

For Scribbright, this might include definitions of AI writing tool, retrieval-augmented generation (RAG), citations in AI answers, grounding, and other terms that explain modern writing, SEO, and AI visibility.

The file can point to answer-ready pages. The pages still need to be worth pointing to.

Retrieval and RAG

Retrieval-augmented generation (RAG) systems retrieve source material before generating answers.

A curated file can help a RAG system identify which pages may be useful sources. The system may still retrieve content through search, crawl data, embeddings, a vector database, or internal indexing, but the guide gives it an explicit list of high-value resources.

That may be especially useful for large or complex sites.

The file tells the system where the good material is. Assuming, of course, that there is good material. The file cannot solve that part, which is inconsiderate but fair.

Embeddings and retrieval systems

Embeddings help AI systems compare text by meaning rather than exact wording.

A site guide does not create embeddings by itself. It can help identify which pages are important enough to retrieve, parse, index, summarize, or embed.

For example, a publisher might highlight its strongest glossary pages, buying guides, comparison pages, and reviews. An AI system could use those links to better understand the site’s main topics and source-worthy content.

The embeddings happen elsewhere. The guide simply improves the selection. The map is not the territory, but it can keep the tourist out of the footer.

Grounding and source support

Grounding ties AI answers to sources, facts, documents, data, or real-world context.

A root-level guide can support grounded answers by pointing systems toward source material that is clear, current, and useful.

Good source candidates usually have:

  • Direct definitions
  • Specific examples
  • Clear claims
  • Useful comparisons
  • Strong internal links
  • Current information
  • Author or publisher context
  • Transparent editorial standards

The file can help AI systems find better sources. It does not guarantee that a model will use them correctly. Source support still has to be checked claim by claim.

AI citations

AI systems may include citations, links, or source references in generated answers.

A curated guide may help citation potential by pointing systems toward pages that deserve attribution, such as definitions, reviews, buying guides, comparison pages, original research, documentation, or editorial standards.

But citation potential is not citation control. The file cannot force an AI system to cite a page. It can only make useful pages easier to identify.

The guide can open the door. The page still has to avoid embarrassing itself once inside.

Share of Model

Share of Model measures how often, prominently, and favorably a brand, site, source, or product appears in AI-generated answers compared with competitors.

A curated guide may support that visibility by clarifying what the site covers and which pages best represent its authority.

For a publisher, this could matter for prompts such as:

  • What are good writing tools for freelancers?
  • What are useful desk tools for writers?
  • What are reliable sources for writing software reviews?
  • What is generative engine optimization?
  • How do AI citations work?

If the file clearly identifies the site’s best resources, it may improve the odds that AI systems understand the site’s relevance.

May is doing real work there. This is not a vending machine.

SEO writing

The file affects SEO writing indirectly.

It is not a replacement for useful pages, clear definitions, strong reviews, accurate comparisons, or well-structured content. It does not make a page credible. It does not repair stale content. It does not create expertise from a list of URLs.

Creating the file can still improve editorial judgment because it forces a site owner to ask:

  • What is this site really about?
  • Which pages best represent the site?
  • Which resources deserve AI attention?
  • Which pages are thin or outdated?
  • Which categories are underdeveloped?
  • Which pages should be improved before they are promoted?

A weak site may struggle to create a useful guide because there is not much worth curating. That is not a technical problem. It is an editorial diagnosis.

Content workflow

A content workflow can include maintenance for the AI-readable guide.

As a site adds glossary pages, product reviews, buying guides, and comparison pages, the file should be reviewed and updated.

A practical workflow might include:

  1. Publish a strategic page.
  2. Decide whether it belongs in the file.
  3. Add the page with a short, useful description.
  4. Remove outdated or weak pages.
  5. Keep section headings clear.
  6. Check links periodically.
  7. Review the file quarterly or after major content changes.

The point is not to list everything. The point is to curate what matters.

Glossary pages

Glossary pages can be strong candidates because they provide direct definitions and structured explanations.

For Scribbright, a useful glossary section might point to pages such as:

The full glossary can also be included as a top-level reference hub.

This helps AI systems understand that the glossary is not a loose pile of definitions. It is a reference library with a strategy. Ideally.

Product reviews

Review sites can use the file to point AI systems toward high-value review content.

A useful review section might include:

  • Best product reviews
  • Main review categories
  • Hands-on review pages
  • Editorial standards
  • Affiliate disclosure
  • Product comparison hubs

For this site, for instance, that would include reviews of writing apps, editing software, keyboards, desk mats, notebooks, pens, focus tools, and productivity resources.

The file should not include every minor review automatically. It should include pages that best represent the site’s expertise, usefulness, and audience fit. Curation is the point. A pile is not a strategy.

Buying guides

A buying guide can be a strong candidate because it summarizes selection criteria, product categories, tradeoffs, and recommendations.

A site guide might highlight buying guides such as:

  • Best writing apps for freelance writers
  • Best keyboards for long writing sessions
  • Best grammar checkers for content marketers
  • Best desk tools for writers
  • Best productivity tools for writers

These are the kinds of pages AI systems may need when answering recommendation prompts.

If the guide is useful, current, and clear, it may deserve inclusion. If it is thin, outdated, or vague, fix the guide first. Do not put a spotlight on the weak floorboard.

Comparison pages

A comparison page can help AI systems understand distinctions between tools, formats, products, or concepts.

Useful examples might include:

  • Writing app vs. word processor
  • Mechanical keyboard vs. low-profile keyboard
  • Grammar checker vs. editing software
  • Gel pen vs. rollerball pen
  • Desk mat vs. mouse pad

Comparisons are useful because AI systems often answer comparative questions. They are also useful for readers who need to make decisions.

The file can point to strong comparisons as reference material. Again, only if the pages actually compare things clearly instead of simply placing two terms near each other and hoping judgment happens.

Benefits

Potential benefits include:

  • Clearer AI-readable site overview
  • Curated links to important content
  • Better support for AI retrieval
  • Cleaner guidance for reference pages
  • Possible support for GEO and AEO work
  • Better context for documentation-heavy sites
  • A useful prompt for internal content audits
  • A simple way to explain site focus

The most practical benefit may be editorial. Creating the file forces a site owner to decide what the site is really about and which pages deserve to represent it.

That exercise can be healthy. Unpleasant, often. But healthy.

Limitations

The file has major limitations.

It is an emerging proposal, not a universally adopted formal standard. Not every AI company, model, crawler, agent, search system, or retrieval tool reads it. Some may ignore it entirely.

Limitations include:

  • No universal adoption
  • No guaranteed AI visibility
  • No guaranteed citations
  • No crawl control
  • No indexing control
  • No training opt-out
  • No substitute for strong content
  • No substitute for schema
  • No substitute for sitemap.xml
  • No substitute for robots.txt

This does not make it useless. It makes it a low-cost, experimental optimization. Low-cost experiments are fine. Just do not put them in the strategy deck wearing a cape.

Common mistakes

Most mistakes come from expecting the file to do too much.

Common mistakes include:

  • Treating it as an official universal standard
  • Assuming major AI systems will read it
  • Using it as a robots.txt replacement
  • Using it as a training opt-out
  • Listing every URL on the site
  • Including thin or outdated pages
  • Writing vague page descriptions
  • Forgetting editorial standards and disclosure pages
  • Failing to update the file
  • Ignoring schema and sitemaps
  • Expecting the file to fix weak content

The biggest mistake is mistaking a map for the destination. The file can point to useful content. It cannot make the content useful.

How to create one

A practical process looks like this:

  1. Define the site’s main purpose.
  2. Identify the primary audience.
  3. List the most important content categories.
  4. Select the pages that best represent each category.
  5. Write a short description for each link.
  6. Use clean Markdown formatting.
  7. Keep the file concise.
  8. Upload it to the root of the domain.
  9. Check that it is publicly accessible.
  10. Review it when major content changes.

For a site like the one you’re on now, the file should likely include the homepage, reviews page, glossary index, major glossary pages, key review categories, buying guides, comparison pages, editorial standards, and affiliate disclosure.

It should not become a dumping ground. That is what old shared drives are for.

A practical checklist

A useful checklist may include:

  • File is named llms.txt
  • File is placed at the domain root
  • File is publicly accessible
  • Content is written in Markdown
  • Site purpose is clear
  • Audience is clear
  • Important pages are curated
  • Descriptions are specific
  • Outdated pages are excluded
  • Glossary or reference hubs are included when relevant
  • Review, buying guide, and comparison hubs are included when relevant
  • Policy or disclosure pages are included when useful
  • The file is reviewed periodically
  • robots.txt, sitemap.xml, and schema remain in place

The final item matters. This file is an addition. It is not a replacement for the boring foundational files that still do real work. The boring things endure.

Who should use one

An AI-readable site guide may be useful for sites with structured, reference-worthy, or complex content.

Good candidates include:

  • Documentation sites
  • Software companies
  • Publishers
  • Glossary sites
  • Review sites
  • Affiliate sites
  • Knowledge bases
  • Research libraries
  • Product information hubs
  • Educational sites

It may be less useful for tiny brochure sites, thin affiliate sites, sites with little original content, or sites that cannot identify which pages matter most.

If choosing the important pages feels impossible, the file may not be the problem. The content strategy may be.

Related tools and concepts

This topic is related to others that include robots.txt, sitemap.xml, schema, retrieval-augmented generation, embeddings, grounding, citations in AI answers, Share of Model, SEO writing, answer engine optimization, generative engine optimization, content workflow, buying guides, comparison pages, product reviews, and affiliate marketing.

Scribbright’s reviews section covers tools, gear, and resources for working writers who want better systems for drafting, editing, researching, reviewing, publishing, and managing their work.

Frequently Asked Questions

What is LLMs.txt?

LLMs.txt is a proposed Markdown file placed at the root of a website to help large language models, AI agents, and retrieval systems find and understand the site’s most important content.

Is it an official standard?

Not in the same way as long-established web standards. It is an emerging proposal and convention, so support may vary across AI systems, crawlers, tools, and platforms.

Does it replace robots.txt?

No. robots.txt is used for crawler access guidance. A Markdown site guide is used to point AI systems toward important content and explain why those pages matter.

Does it stop AI companies from training on my content?

No. The file is not a training opt-out mechanism. Site owners need separate technical, legal, or crawl-control methods if they want to restrict access or communicate usage limits.

What should go in the file?

Useful items include a short site description, audience context, core pages, glossary or documentation links, product review hubs, buying guides, comparison pages, editorial standards, and disclosure pages.

Will it make AI systems cite my site?

No. It may help AI systems find citation-worthy pages, but it does not guarantee citations, inclusion, traffic, recommendations, or visibility in generated answers.

Key takeaways

  • LLMs.txt is a proposed Markdown file that gives AI systems a curated guide to a site’s most important content.
  • It usually lives at the root of a domain, such as https://example.com/llms.txt.
  • It does not replace robots.txt, sitemap.xml, schema, noindex controls, or AI training opt-out methods.
  • It may support AI retrieval, GEO, AEO, grounding, citations, and Share of Model, but it guarantees none of them.
  • The best version is concise, specific, curated, publicly accessible, and maintained as the site changes.

Browse more definitions in the Scribbright glossary.

Scroll to Top