
Prompt research is the practice of identifying the actual questions buyers ask AI systems about your category, then using them as the measurement baseline and content brief for AI visibility work. It replaces keyword volume as the demand signal, because conversational prompts are longer, unmeasured by traditional tools, and expanded further by query fan-out.
Keyword tools measure what people type into a search box. They do not measure what people ask an assistant, and the two have diverged in three ways.
You are competing for questions nobody typed. Fan-out means a page can be assessed against sub-questions with no search volume attached, which is precisely the surface a keyword-only strategy cannot see.
This has an uncomfortable budgeting implication. Keyword volume let teams forecast: a term with a known monthly volume and an estimated click-through rate produced a business case. Prompt research offers no equivalent, which means AI visibility work has to be justified on presence and conversion quality rather than on projected traffic. Teams that try to force the old forecasting model onto the new channel usually end up either overstating the case or abandoning it.
| Search query | Prompt | |
| Length | One to four words | Ten to forty words |
| Form | Keyword fragment | Full question or instruction |
| Context | None carried | Often carries constraints, budget, situation |
| Follow-up | New search | Conversational refinement |
| Measurable | Yes, by volume tools | No, only by testing |
| Competitive field | Ten results | Three to fifteen cited sources |
The context point is the one with strategic weight. A prompt frequently contains the qualifying information a salesperson would ask for: company size, budget, timeline, constraints. That means a well-matched answer is closer to a qualified lead than a ranking ever was, which is consistent with Semrush finding AI search visitors convert at 4.4 times the rate of traditional organic visitors.
The follow-up behavior matters too. A search session that fails produces a new query; a prompt that fails produces a refinement in the same conversation, with the earlier context retained. Practically, this means a brand named early in a conversation stays present as the buyer narrows down, which raises the value of appearing in the broad opening question rather than only in the specific final one.
Prompts cannot be pulled from a tool, so they have to be collected. Five sources produce the language buyers actually use.
A sixth source is worth adding once you are running: the prompts themselves. Testing your set surfaces adjacent questions, competitor framings, and category vocabulary you had not accounted for, which feed the next revision.
Collect verbatim, not paraphrased. The value is in the exact phrasing, including the imprecise terms buyers use before they learn your category’s vocabulary. Work on search intent is the analytical layer on top of this collection, not a substitute for it. Sales and support teams are usually willing collaborators here, since the exercise surfaces questions they have been answering repeatedly without anyone documenting them.
Different prompt types are won by different work, which is why a mixed set is essential.
| Type | Example shape | What wins it | Owner |
| Category | What is X and how does it work | Clear explanatory content | Content |
| Comparative | Best X for Y, alternatives to Z | Third-party roundups and reviews | PR and partnerships |
| Problem-led | Why is my X doing Y | Demonstrated expertise, specific answers | Content and support |
| Navigational | What does company X do | Entity clarity and consistency | Brand and entity work |
Balance the set across the buying journey as well as across types. Early-stage questions establish presence before a shortlist forms, and late-stage questions decide it. A set covering only one end measures a fraction of the surface that matters.
A set weighted entirely toward category prompts will show early wins and no commercial movement, because the money sits in comparative and problem-led questions. A set weighted entirely toward comparative prompts will show nothing for two quarters, because those are the slowest to move.
Navigational prompts deserve more attention than their volume suggests. Asking an assistant what your company does is the fastest available diagnostic for entity clarity, and the answer determines how you are described in every other prompt where you appear. A brand that cannot be described correctly by name will not be described correctly in a comparison either.
Thirty to fifty prompts is the right size. Enough to be representative, small enough to test repeatedly by hand.
Freezing the set is the step teams skip and later regret. A prompt reworded between quarters produces a change in results that cannot be attributed to anything.
Record more than presence on each run. At minimum: whether you appeared, in what position within the answer, how you were described, which domains were cited as evidence, and which competitors appeared alongside you. The cited domain column is the one most often omitted and the one that most often changes what you work on next, since it identifies the specific sources supplying the answer.
Near-identical prompts return different brand sets on different runs. This is a property of the systems, not an error in your testing.
Handle it with method rather than by ignoring it:
Semrush found that 45% of marketing leaders cannot accurately measure brand visibility inside AI-generated answers, and only 9% have tools covering all relevant metrics. A frozen prompt set tested by hand outperforms most tooling precisely because the method is controlled.
Volatility is also informative rather than merely inconvenient. A brand appearing in one run of three has weaker corroboration than one appearing in all three, even though a single test would record both as present. Tracking that fraction over time gives you a sensitivity measure: rising consistency on the same prompts usually precedes a rise in overall presence, which makes it an early indicator worth reporting.
Not every prompt is worth competing for. Score each on three dimensions and work from the top.
| Dimension | Question | Weight |
| Buying proximity | How close is this question to a purchase decision? | High |
| Winnability | Is the current answer generic, or dominated by an entrenched incumbent? | High |
| Effort | Does winning it require content, corroboration, or both? | Medium |
Category concentration sets expectations here. Semrush found the three most visible brands hold 82.9% of category visibility in News and Media and 76.9% in Consumer Electronics, against 42.2% in Industrial and 41.4% in Finance. In concentrated categories, the winnable prompts are narrow and specific rather than category-level.
Look for prompts where the current answer is vague. An assistant giving a generic response to a specific question is an opening. An assistant naming three entrenched brands confidently is not, at least not this quarter.
Two further prioritization signals are worth logging. Prompts where the cited evidence comes from sources you could plausibly appear in are more winnable than those citing sources you cannot reach. And prompts where competitors appear inconsistently across runs indicate an unsettled answer, which is a better target than one where the same three names appear every time.
The prompt set is a work queue once you know why you are losing each one.
For every prompt where a competitor appears and you do not, classify the gap:
Each gap type produces a different brief. A coverage gap needs a page. A citability gap needs an hour of restructuring on an existing page. A corroboration gap needs work nobody in the content team can do, which is why building an AI content strategy from a prompt map produces better allocation than building one from a keyword list.
The distribution of gap types is itself a useful finding. A set dominated by citability gaps means your content is adequate and badly structured, which is the cheapest situation to be in. A set dominated by corroboration gaps means the content team cannot solve the problem regardless of budget, and continuing to commission articles will not change the outcome. Diagnosing this before allocating a quarter’s spend is the main practical return on building the set at all.
Google adds a boundary worth respecting: it warns that creating separate content for every query variation, including fan-out queries, risks its scaled content abuse policy. A prompt map is a guide to what to cover properly, not a list of pages to generate.
A prompt set decays. Categories develop new vocabulary, products change, and competitors reposition.
| Activity | Cadence |
| Run the full set | Monthly |
| Review competitor set for new entrants | Quarterly |
| Add prompts from new sales and support language | Quarterly |
| Retire prompts no longer reflecting real questions | Every six months |
| Rebuild the set from scratch | Annually, or after any repositioning |
Assign an owner. A prompt set with no named custodian is run enthusiastically for two quarters and then quietly abandoned, usually at the point where the data would have become most useful for showing trend.
Keep retired prompts in the record rather than deleting them. Historical comparability is the reason the set exists, and a full rebuild should be a deliberate reset with a new baseline rather than a drift.
Two triggers justify an unscheduled rebuild. A significant repositioning changes which questions you should be competing for, which makes an old set measure the wrong thing accurately. And a major platform change, such as a new default model or a shift in how a system fans out queries, can move results across the whole set at once, in which case the useful record is the note explaining what changed rather than the movement itself.
A frozen, well-built prompt set is the single most useful asset in AI visibility work. It is the baseline, the work queue, and the reporting line at once, which is why a generative engine optimization program should produce one before proposing anything, and why the composition of that set is worth reviewing together when you get in touch.
Can I get prompt volume data like keyword volume?
No. Prompts entered into AI assistants are not reported by any platform, so no tool has access to them. Any vendor offering prompt volume figures is modeling from search data rather than measuring actual prompts, and those estimates should be treated accordingly.
How many prompts should I track?
Thirty to fifty is the practical range. That is large enough to be representative of a category and small enough to test three times each across multiple platforms by hand, which is necessary because near-identical prompts return different results on different runs.
How is prompt research different from keyword research?
Keyword research measures what people type into a search box, using reported volume. Prompt research collects what people actually ask assistants, which is longer, question-shaped, carries context, and is unmeasurable. Query fan-out compounds the difference, since systems generate their own sub-questions that no keyword tool tracks.
Should I create a page for every prompt?
No. Google warns that producing separate content for every query variation, including fan-out queries, risks its scaled content abuse policy, and states that a high quantity of pages does not make a site more relevant. Use the prompt map to decide what to cover properly, not how many pages to publish.
How often should I re-test my prompt set?
Monthly for the full set, with the wording frozen so results stay comparable. Add new prompts quarterly as sales and support language evolves, and rebuild the set from scratch annually or after any significant repositioning.