Complete Guide

Present Is Not Useful: What a Machine Receives From Your Pages, and Why We Rebuilt Curator Around It

A page can be listed, described and served to a machine and still say nothing useful. What we observed on real sites, what Curator 2.2 does about it, and the Museum Guide that explains it.

9 min read 1,811 words Updated Oct 2026

A page can be listed in an llms.txt, carry a description and be served to a machine as Markdown, and still say nothing useful, because what was served came from the wrong source, dragged the site’s furniture in with it, or lost its structure on the way. We saw all three on real sites. LLMs.txt Curator 2.2 builds each page’s Markdown from a snapshot of the page a visitor sees, shows you what a machine receives page by page, and names the faults instead of scoring them.

2,075 automated converter checks pass on the released 2.2.0 build, covering tables, lists, code and nested structures that the previous converter flattened LLMs.txt Curator 2.2.0 release record, SEO Strategy Ltd, October 2026
1,937 characters between U+0080 and U+FFFF were damaged by the whitespace bug fixed in 2.2.0, including Persian, Arabic, Hebrew, Hindi and some accented Latin letters LLMs.txt Curator 2.2.0 release record, reported by Saeid Afshari, 4 October 2026
262 checks of the read-only agent layer, including the permission matrix, enumeration and hostile input, pass on the released build; no ability can write LLMs.txt Curator 2.2.0 release record, SEO Strategy Ltd, October 2026

This is the second research-led piece in our LLM Optimisation series. The first, Ranked but Not Recommended, was about search. This one is about your own site, and the plugin we maintain for it, LLMs.txt Curator, now at version 2.2.

What does “present” mean, and why is it not enough?

An llms.txt file lists the pages you want a machine to read, each with a short description. Version 1 of Curator measured whether those descriptions existed. It could tell you that 45 of 48 pages had one, and that felt like the right thing to measure.

Then we ran version 2.0 across a real product catalogue and watched it report excellent coverage on a file that was close to useless. Dozens of product descriptions opened with the same line of shop furniture, a call to action repeated across the whole catalogue, because that was the first text each page actually published. Every description was present. Almost none was useful.

That is the whole argument in one observation. Presence is a count: is there a description, is there a Markdown version, is the page in the file. Usefulness is a judgement: does what the machine receives say what the page says? A count cannot answer the second question, and a tool that only reports counts will tell you everything is fine while it is not.

Where does the gap come from?

Three places, and we found all three on live sites rather than in theory.

  • The stored post is not the page. WordPress keeps the text an author typed. A page builder or a block template adds pricing tables, service grids and calls to action at render time. A Markdown version built from the stored post leaves all of that out, so the machine gets the skeleton and the visitor gets the page.
  • The page is more than its content. Convert the rendered page instead and the opposite happens: menus, footers, share buttons, comment forms and “more posts” lists leak into the Markdown. A page about retrofit assessments ends up half about the site’s navigation.
  • Converters lose structure. A regular-expression converter flattens a pricing table into one run-on line, so a machine reads “ServicePriceAudit” where the page said three separate things. Numbered lists lose their numbers; nested lists go flat; code loses its indentation.

And one more that is not about structure at all. A whitespace pattern in the plugin, run in byte mode, matched a byte that sits inside letters in most non-Latin scripts. On a Persian site it turned a word into a question mark and an invalid file. The description was present. The description was wrong. Nothing in a coverage count could have seen it. Saeid Afshari of kabook.ir did, reported it, and is credited in the changelog.

The Museum Guide

The physical picture we use for this is a museum.

Your website is the museum. Visitors walk the rooms and look at the exhibits, which are your pages. The llms.txt is the printed guide at the entrance: a short list of the rooms worth seeing and a line about each. A machine arriving at your site is a visitor who reads the guide first.

Version 2.0 of Curator was about the guide. Which rooms are listed, who decided, what the line about each one says, and whether the guide at the entrance is the one you approved or an old print run somebody forgot to replace.

Version 2.2 is about the labels on the exhibits. When a machine follows the guide into a room, the Markdown version of the page is the label it reads. If the label was written from the curator’s notes rather than from the exhibit itself, it describes the painting that was meant to hang there, not the one that does. If the label has the fire-exit sign and the gift-shop opening hours printed across it, the visitor cannot tell which words are about the painting. If the label has been transcribed by someone who cannot read the alphabet it was written in, it says something, and what it says is wrong.

The job of a good museum guide is not to write new labels. It is to walk the rooms, read each label standing in front of the exhibit, and say plainly where they disagree. That is what the Markdown view and the Pages checks in 2.2 do.

What did we actually observe?

These are the observations the release was built from. They are specific to the sites they were seen on, and we do not generalise from them to anyone else’s site.

What we observed
Dozens of product descriptions beginning with the same call to action
Where: a live product catalogue during 2.0 testing. What it meant for a machine: present, not useful. Coverage was reported as excellent.
A pricing table flattened into one run-on word
Where: a description suggested “from page content”. What it meant for a machine: three facts read as one unreadable token.
Persian text published with a question mark in place of a letter, and an invalid file
Where: kabook.ir, reported 4 October 2026. What it meant for a machine: wrong text served, and because WordPress refused to store the invalid file, the stored copy and the file on disk disagreed.
A synthetic site whose theme listed all 10,000 pages in its navigation
Where: the release test bed. What it meant for a machine: every page rendered at 2.2 MB; 18 were refused as too large and the stored content was used instead.

The last row is a test artefact, not a client. It is in the table because it shows the failure mode cleanly: a page can be fine for a visitor and unusable for a machine purely because of what surrounds the content.

What does Curator 2.2 do about it?

It builds the label from the exhibit, lets you read the label, and says where it disagrees.

  • Snapshot, not stored post. With Markdown pages on, each page in your llms.txt is fetched from your own site as a logged-out visitor, in the background. The Markdown is built from the page’s main content area, so page-builder and template content is included and the surrounding furniture is left out.
  • A converter that keeps structure. Tables stay tables, lists keep their numbers and nesting, code keeps its indentation. One converter writes every Markdown address and llms-full.txt, so they cannot disagree with each other.
  • A Markdown view per page. What the machine receives, laid out for reading, with how it was made: snapshot or stored content, which part of the page was used, what was removed, what was kept.
  • Checks that name the problem. Pages built from stored content because a snapshot failed, and why. Pages with very little text. Unrendered shortcodes. Leftover markup. Repeated text. Uneven tables. Menus that leaked in. Observed facts, not a score.
  • Never stale. A snapshot is retired the moment its page changes, and when a post is unpublished or made private every page that linked to it is retired too, so it leaves lists of posts straight away.
  • Byte-safe text. One whitespace pattern that treats non-Latin letters as letters, covered by a test of every character from U+0080 to U+FFFF.

Two things it deliberately does not do. It does not write or rewrite your text; the honest answer to a bad label is a human, not a model. And it does not score pages. A score would invite the same mistake coverage invited: a number that looks like a verdict.

What does this have to do with agents?

Once the label on each exhibit is trustworthy, you can let someone else read it. Version 2.2 opens Curator to signed-in tools and agents through WordPress’s own Abilities API, read-only: list what the llms.txt publishes, get a page’s Markdown exactly as its .md address serves it, and ask what changed since last time. With WordPress’s MCP Adapter installed, an AI client such as Claude can use the same four abilities.

The rules are strict because the alternative is not worth having. The abilities answer only from the record of what you published, never from your working curation. Anything not published gets one identical “not available” answer, so nothing hidden can be enumerated. No ability can change curation, settings or what is published. Anonymous requests are refused. In the museum, the audio guide reads the labels; it does not get a key to the storeroom.

Evidence ledger
What the release record demonstrated
2,075 converter checks, 136 representation checks, 262 agent-layer checks and 47 multibyte checks pass on the unpacked release ZIP, against WordPress 7.1.2 on PHP 8.4. Checks across 10,000 entries run in 2.75 seconds against a 10-second budget. The non-Latin defect is reproduced on 2.1.0, which fails 45 of the 47 multibyte checks.
What it did not say
Nothing here measures whether any AI system reads a given site’s llms.txt or Markdown, or what it does with it. Google Search ignores the file. No ranking, citation or traffic effect is claimed for any of this work.
What we infer
That counts of presence will keep passing files that machines cannot use, and that the only reliable check is to look at what a machine receives, page by page, built from the page a visitor sees. This is an inference from the sites we observed, not a general law.
What someone observed
A Persian site owner reported damaged text and a broken Settings link on 4 October 2026. A catalogue site showed excellent description coverage with near-identical descriptions. Both are the observations that set the direction of 2.2.

What should you do first?

  • Look before you count. Open three pages’ Markdown versions, the ones you would least like a machine to misread, and read them standing in front of the page. If you use Curator, that is the Markdown view in Your llms.txt.
  • Check the furniture. If your menu, footer or related-posts list appears in the Markdown, the machine is reading the building, not the exhibit.
  • Check the structure. Pricing and comparison tables are where the most meaning lives and where converters lose the most. If a table reads as one line, fix the source or the converter before anything else.
  • If your site is not in a Latin script, check a page by eye. A pattern bug can pass every automated test written by someone whose alphabet it does not touch.
  • Then curate the guide. The entrance guide still matters: which pages, in what order, with what line. Our llms.txt guide covers that, and Curator is free.

Where this sits in the wider work: being present is the bottom of the AI Discovery Stack, and most of the Stack is about what happens after a machine has found you. But a machine that reads the wrong label learns the wrong thing, and nothing further up the Stack corrects that. Strong brands rank, get cited, remembered, recommended and dominate; a brand whose own pages describe it wrongly to machines is handing that work to someone else.

Key Definitions

Representation
The version of a page a machine receives: its line in llms.txt and, with Markdown pages on, its Markdown at the page’s address plus index.md.
Snapshot
Curator’s copy of a page as a logged-out visitor receives it, fetched from the site itself in the background and used to build the page’s Markdown.
Present versus useful
Present: a description or Markdown version exists. Useful: what it says matches what the page says. Counts measure the first; only looking measures the second.
Abilities API
WordPress’s own way for plugins to describe capabilities that tools and agents can call. Curator exposes eight read-only abilities through it.
Museum Guide
Our teaching picture for this: the llms.txt is the printed guide at the entrance, each page’s Markdown is the label on an exhibit, and the job is to read every label standing in front of its exhibit.

Frequently Asked Questions

What does “present is not useful” mean?

A description or a Markdown version can exist for every page and still be wrong: built from the stored post rather than the rendered page, polluted by menus and footers, flattened by a poor converter, or damaged by a text-handling bug. Coverage counts measure existence. Usefulness is whether what the machine receives matches what the page says, and you have to look to know.

Does any of this affect Google rankings?

No. Google Search ignores llms.txt, and nothing in this piece claims a ranking, citation or traffic effect. The work is about controlling and verifying what your own site exposes to systems that do read it, such as Perplexity, coding agents and signed-in tools.

Why build the Markdown from a snapshot rather than the stored post?

Because the stored post is not the page. Page builders and block templates add content at render time that never reaches the database, so a Markdown version built from the stored post leaves it out. A snapshot taken as a logged-out visitor captures what the page actually shows, and the extractor then keeps the main content and drops the surrounding furniture.

What was the non-Latin bug?

A whitespace pattern run in byte mode matched a byte that sits inside letters in most non-Latin scripts, so Persian, Arabic, Hebrew, Hindi and some accented Latin characters were damaged in llms.txt and llms-full.txt. It was reported by Saeid Afshari of kabook.ir on 4 October 2026, fixed in 2.2.0, and is now covered by a test of every character from U+0080 to U+FFFF.

Can an AI agent change my llms.txt through Curator?

No. The eight abilities are read-only and answer only from the record of what you published. Four are offered to MCP clients. Unpublished, draft, private and excluded pages all receive the same not-available answer, and anonymous requests are refused.

Is the Museum Guide a real feature?

It is a teaching picture, in the same family as the Lift and the Research Desk, not a product name. The features it explains are the Markdown view, the Pages checks and the read-only abilities in Curator 2.2.

Sean Mullins

Founder of SEO Strategy Ltd with 20+ years in SEO, web development and digital marketing. Specialising in healthcare IT, legal services and SaaS — from technical audits to AI-assisted development.

Ready to improve your search visibility?

Book a free 30-minute consultation and let's discuss your SEO strategy.

Get in Touch