← The journal
Search Method

How People and LLMs Search for AI Infrastructure Answers

What keyword data and search docs show about how people and AI answer engines look for AI infrastructure answers.

Swfte Journal / Search Method

We built a section of this site made of how-to guides, and we wanted to know what shape those guides should take before we wrote 26 of them. This post is the method: which data we looked at, what it showed, what it could not show, and the decisions that followed. It makes no claim about how any answer engine ranks pages, because we have not measured that and the vendors do not publish it. Where we are guessing, we say so.

The short version: people and machines ask the same questions in different forms. People type a question and scan for a numbered list. An answer engine breaks the question into several narrower searches, fetches the pages that come back, and quotes the one that states the answer plainly. A page that serves the first reader well mostly serves the second. The differences are small, specific and cheap to fix, and they are mostly about verifiability.

The two readers

A person with a task. They have a goal ("run a model on my laptop"), a constraint ("16 GB of memory, no GPU") and a worry ("will this break something"). They search in question form, glance at the top results, open two or three, and look for steps they can follow. They leave a page that opens with a history of the technology. They come back to a page that told them how long it would take and what they would see at each step.

An answer engine or a coding agent. It receives a prompt, decides what it needs to look up, and issues several searches rather than one. Google describes this for its own AI features: they "may use a 'query fan-out' technique", which it defines as "issuing multiple related searches across subtopics and data sources". The same page, last updated 10 December 2025 when we read it, says there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary", that you "don't need to create new machine readable files, AI text files, or markup", and that a page needs to be "indexed and eligible to be shown in Google Search with a snippet". It also says "indexing and serving isn't guaranteed".

That last point shapes everything below. If the plain requirement is a page that is indexable and useful, then the work is in the page itself, not in a new file format. We still publish machine-readable twins of our guides (more on that later), but we treat them as a convenience for tools that fetch pages, not as a ranking lever.

What we looked at

We used three sources, and recorded the date for each.

  1. A DataForSEO keyword pull dated 6 October 2026. It holds 6,872 rows for the AI infrastructure and governance markets we cover, each with a modelled monthly volume, a cluster, and SERP feature flags. DataForSEO volumes are modelled estimates, not measured traffic, and we treat them as a map of relative demand, not a forecast.
  2. WebSearch checks of the head query and one or two variants for each guide topic, run on 6 October 2026. We recorded the page types and domains in the top results and any People Also Ask or related-search text the tool returned.
  3. The documentation of the platforms involved: Google's page on AI features, OpenAI's crawler documentation and the llms.txt proposal.

We did not read any chat logs or interview anyone. We have no measurement of what an answer engine actually did with our pages, because none of them gave us that. Everything here is about the demand side and the public rules.

What the keyword data showed

Question-form queries are a small share of the volume. Of the 6,872 rows, 507 start with a question word (what, how, why, when, which, can, does, is, are, should, who). Together they carry 143,736 of the 2,709,933 baseline monthly searches in the file, about 5 per cent. People looking for AI infrastructure answers mostly type nouns ("ollama docker", "litellm vs openrouter"), not sentences. A question-shaped H1 helps because it matches the sentence an answer engine writes, not because most searches are sentences.

Definitions dominate the question volume. In the governance cluster, 86 question keywords carry 81,350 searches, and the top five are definitional: "what is agentic ai" (33,100), "what is an ai agent" (18,100), "what are ai agents" (4,400), then "how to build an ai agent" (2,900) as the first how-to. In AI security the top question keywords are "what is shadow ai" (1,900) and "what is prompt injection" (1,300). In compliance, "what is the eu ai act" (320) and "what is iso 42001" (260). The first step of a how-to is often a definition in disguise, so each guide opens with a short answer and defines terms in line.

How-to phrasing is rarer than it feels. Across the file, 133 rows begin "how to" or "how do", and 103 of them were tagged by the pipeline as wanting a how-to page. The rest are API-key and pricing questions. Phrases people often assume are common barely appear: we found no keyword containing "step by step", none containing "error", and 12 containing "tutorial" (390 searches in total). The universe was seeded from product and topic terms, so an absence is not proof nobody searches them. It does mean a keyword tool alone will not tell you the wording of the sub-queries an answer engine generates.

Commercial modifiers carry far more volume. 441 keywords contain "pricing" (87,103 searches), 229 contain "vs" (43,414), and 61 contain "github" (17,424). So every guide gets a short section on cost, a comparison where one tool is the obvious alternative, and a link to the repository or docs for the tool it uses.

SERP features are everywhere in this market. 3,458 of the 6,872 rows carry a People Also Ask flag and 3,546 carry an AI Overview flag. Half the market shows an answer box above the links. Ranking for a head term is worth less than it was, and being the source an answer box quotes is worth more.

What the SERP checks showed

For each guide topic we ran the head query and one variant through WebSearch. The pattern was consistent across the six topics we had finished at the time of writing (RAG, company brain, gateways, migration, cost reduction, agent monitoring):

  • The results were almost all vendor and agency blog posts, not official documentation. Many titles carry a year and a count or a percentage ("10 proven strategies", "cut your API spend by 70-90%").
  • Several of the percentage claims had no source. In one case a blog post's figure for prompt caching disagreed with the vendor's own documentation, which we read the same day. A page that states the vendor's figure with a date and a link is checkable. The blog post is not.
  • The tool returned no People Also Ask or related-search text for any query. We wrote "none returned" in our notes rather than inventing questions. The FAQs on each guide are built from the question keywords in the DataForSEO file instead.

The fan-out sub-queries we designed for

Because we cannot observe an engine's sub-queries, we wrote down the ones a careful person or agent would plausibly issue for a task like "set up an LLM gateway", and made sure the page answers each:

Sub-query shapeExampleWhere the page answers it
Steps"how to set up an llm gateway step by step 2026"Numbered steps, each with a one-line outcome
Comparison"litellm vs openrouter"A short comparison table with links to full pages
Requirements"llm gateway requirements"Prerequisites and a hardware or account list
Cost"llm gateway pricing"A cost line that names what is free and what is not
Licence"litellm licence"A sources list that includes the licence page
Tutorial"litellm proxy tutorial github"Links to the official repository and docs
Error"litellm proxy 401 error"A troubleshooting table with the error text

These are our guesses, recorded as the relatedQueries field of each guide, which the topic map uses to decide which page owns which question. They are not observations.

What we changed because of this

  1. Answer first. Each guide opens with a short answer of 40 to 90 words that can be quoted on its own.
  2. One H1 phrased as the question, one primary intent per page, and a clean, flat URL such as /how-to-run-llms-locally that never changes.
  3. Numbered steps with a one-line outcome, copyable commands, and what the terminal should show.
  4. A last-verified date and a sources list on every guide, with the official page each command was checked against. If something could not be verified we left it out.
  5. A troubleshooting table using real error text, because "X error" is a standard sub-query.
  6. A markdown copy and a JSON index at /how-to/<slug>.md and /how-to/index.json. The llms.txt proposal, which Jeremy Howard published in September 2024, recommends markdown copies at the same URL with a .md suffix. Its own page says labs publish the file for their developer docs and does not claim that search engines read it, so we do not assume a benefit.
  7. We left crawler rules to robots.txt. OpenAI documents three agents: OAI-SearchBot for search, GPTBot for training, and ChatGPT-User for visits a user triggers, noting that robots.txt rules "may not apply" to the last. Each can be set separately. Blocking the search bot means "not shown in ChatGPT search answers". That is a decision for your site, not something a page can influence.

Limits of this study

  • Modelled volumes are not traffic.
  • The seed list shapes what the keyword file can contain.
  • WebSearch returns summaries, not the live results page, and showed no People Also Ask text, so we could not study it.
  • We observed the demand side only. Whether any of these changes helps a page get cited is a hypothesis we will test by looking at referrals and at what the engines say when asked.

Do this for your own site

  1. Pull your question keywords and count them honestly. Note which are definitions and which are tasks.
  2. Search each head query and write down the page types in the top ten and the dates. Note what is missing: ours was official-source verification.
  3. List the sub-queries for the task, cheap ones first, and check that your page has a section for each.
  4. Put the answer first, date it, and cite the official source for every command.
  5. Publish a markdown copy if it costs you nothing. Do not expect it to replace the work above.

Our how-to guides, listed on the how-to hub, follow this structure. If you want to see it in practice, start with how to run LLMs locally or how to validate your AI. For the other half of the argument, see how to write pages AI answer engines cite.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.