← The journal
Search Method

How to Write Pages That AI Answer Engines Cite

A page structure for being cited by AI answer engines: answer first, numbered steps, dated sources and clean URLs.

Swfte Journal / Search Method

Nobody outside the answer engines can tell you what makes them cite one page over another. The vendors publish rules about eligibility, not ranking, and every "secret" you read about in a blog post is somebody's experiment on a small sample. What we can offer is narrower and more useful: a list of page properties that are cheap to get right, that do not depend on guesses about an algorithm, and that make a page better for a human reader at the same time. This is the structure behind our how-to guides, and the reasoning for each piece.

For the evidence on how people and engines search, see the companion post on how people and LLMs search for AI infrastructure answers. This one is about the page.

Start from what the vendors say

Google's page on its AI features, which we read on 6 October 2026 (last updated 10 December 2025), makes four statements worth building around:

  • "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."
  • You "don't need to create new machine readable files, AI text files, or markup to appear in these features."
  • To be a supporting link, a page must be "indexed and eligible to be shown in Google Search with a snippet."
  • Structured data should match the visible text, and "there's also no special schema.org structured data that you need to add."

That is a short list, and it puts the weight in the right place: the page has to be crawlable, indexable, and worth quoting. OpenAI's crawler documentation draws a similar line from the other side. It lists OAI-SearchBot ("for search"), GPTBot (content that may be used for training) and ChatGPT-User (visits triggered by a user in ChatGPT). Each can be allowed or blocked separately in robots.txt, and blocking the search bot means your site "will not be shown in ChatGPT search answers, though can still appear as navigational links." The user-triggered agent is described as one where robots.txt rules "may not apply".

Practical consequence one: check your robots.txt and your CDN's bot rules before you touch a single heading. A page that an engine's crawler cannot fetch cannot be cited, whatever it says.

Practical consequence two: serve the answer in the HTML you send. Many fetchers and coding agents read the response body and do not run scripts. If the steps appear only after JavaScript runs, some readers will see an empty page. Our guides are server-rendered, with the commands in the static markup and the copy button as an enhancement.

The structure, element by element

1. One question, one page

Each page answers one question and the title says which. Our flagship is /how-to-create-your-own-local-model, not a section buried in a broader piece. The H1 is the question as people type it. A page that tries to answer three questions gives an engine three weak matches instead of one strong one.

Keep the URL flat, readable and permanent. Do not put a date or a version number in it. When a guide changes, change the page and the last-verified date.

2. The answer in the first screen

State the answer in 40 to 90 words before anything else, and write it so it still makes sense when quoted alone: the outcome, the method in a breath, the main caveat. If someone lifts only that paragraph, they should not mislead their reader.

Do not open with a paragraph about how fast AI is moving. It gives a fetcher nothing to quote and a person nothing to do.

3. Prerequisites and estimates before steps

List what the reader needs (operating system, memory, accounts, tools), how long it will take, and what it costs. These are the follow-up questions every task page gets ("minimum RAM for X", "is it free"). Say them plainly. Label them as your estimates. Do not invent figures: if you cannot derive a number from a source or from arithmetic you show, say "check current pricing" and show how to work it out.

4. Numbered steps with one-line outcomes

A step is a verb and an outcome: "Install Ollama. You end up with the ollama command on your path." Within each step, give the command in a copyable block, then what the reader should see. Numbered steps are easy for a person to scan and easy for a machine to lift as an ordered list. Keep the sequence honest: if step four depends on a choice made in step two, say so in step four.

5. Exact commands and versions, each checked

This is where most competing pages are weakest, and where a careful page is easiest to trust. For every command, flag, package name and version:

  • Read it on the official documentation or the project's repository, on the day you publish.
  • If you cannot verify it, leave it out.
  • Say which version the docs showed when you checked, where the version matters.
  • Link the source next to the step.

Copying a command from another tutorial copies its mistakes. In our own checks of the top results for several topics, popular posts repeated figures (for example a caching discount) that disagreed with the vendor's own documentation read on the same day. A page that states the vendor's figure with a date and a link is checkable, and being checkable is the only durable advantage a small publisher has.

6. A dated "last verified" line

Show a visible date and mark it up (<time datetime="2026-10-06">). Dates in dateModified and a visible line tell a reader, and anything that reads the page, how much weight to give a command. Update the date only when you have actually re-checked the commands, not when you fix a typo. A fresh date on stale content is worse than an old date on content that still works.

7. Troubleshooting with the real error text

"X error" is a standard follow-up search. Give a table with three columns: what you see, the likely cause, the fix. Use the exact error string where you have verified it, so a person who pastes the error into a search box lands on your page.

8. A short checklist to confirm it worked

Five or six checkable items: "the endpoint answers on port 8080", "the held-out set scores above the baseline". This turns a vague "done" into something a person, or a coding agent running your steps, can test.

9. A short FAQ built from real questions

Take the questions from your keyword data and from the question boxes in search results, not from imagination. Put the answer first, in two or three sentences. Do not stuff the FAQ with questions nobody asks.

10. Sources, with what each one confirmed

End with a list of the sources you used and one line on what each one covers. It helps a reader audit your work, and it is the section most likely to be missing on a competitor's page.

Structured data: do it, but know what it buys you

We add HowTo, FAQPage, BreadcrumbList and Article JSON-LD to every guide. We do it because it describes the page accurately to anything that reads it, not because we expect a rich result. Google's own statement is that no special schema is needed for its AI features, and in August 2023 it announced that FAQ rich results would be shown mainly for well-known government and health sites and that How-To rich results would be limited to desktop (and later it stopped showing them). Search coverage of FAQ results has changed again since, so check Google's current structured data documentation before you promise anyone a snippet.

The rule that matters is the one Google states: the markup must match the visible text. Generate the JSON-LD from the same data that renders the page, so the two cannot drift apart. We do exactly that: each guide is one typed record, and the page, the schema, the markdown copy and the index are all produced from it.

Machine-readable twins

We publish a markdown copy of each guide at /how-to/<slug>.md and a JSON catalogue at /how-to/index.json, and list the guides in llms.txt. The llms.txt proposal, by Jeremy Howard, asks sites to provide a markdown file of background and links, and recommends clean markdown versions of pages at the same URL with .md added. Its own page says labs publish the file for their developer documentation, and it does not claim that search engines use it. We treat these files as a courtesy to tools that fetch pages and a cheap, regenerated-from-data output. We do not rely on them.

A pre-publish checklist

  1. Is the page indexable (no stray noindex, canonical points to itself, not blocked in robots.txt)?
  2. Does the H1 match the question, and is there only one H1?
  3. Is there a quotable answer in the first screen?
  4. Does every command have a source and a date?
  5. Does the page say what the reader will see after each command?
  6. Is there a troubleshooting table with real error text?
  7. Do the JSON-LD and the visible text say the same thing?
  8. Does the page link out to the official docs it relied on, and to your own related pages?
  9. Have you removed every figure you cannot source?

What we cannot tell you

We cannot tell you that following this gets you cited. We have no access to the engines' citation decisions, and the study we ran on the search side shows only demand and public rules. The structure above is justified on its own: it is how we would want to be taught a task. If it also earns citations, we will find out from referrals and from asking the engines, and we will say what we saw. To see the format in practice, open how to build a RAG system or how to set up an LLM gateway.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.