Playbook
The GEO Visibility Sprint
Four weeks. Week one you write the prompt set you want to be named in. Week two you run those prompts across the assistants and log what they actually say. Week three you publish the specific pages that fill the gaps you found. Week four you re-run the identical prompts and compare.
Last reviewed 27 August 2026
- Outcome
- A measured baseline of what AI assistants say about your category, and published content aimed at the specific gaps
- Timeframe
- 4 weeks, repeated quarterly
The steps
Run it in this order.
01
Write 40 prompts a real buyer would actually type
Not keywords. Full sentences, in the shape people ask assistants questions. Mix category prompts, comparison prompts, problem prompts, and constraint prompts such as the cheapest option for a team of three. Pull the wording from your own sales calls and support inbox rather than inventing it, because invented prompts measure a market that does not exist.
02
Freeze the prompt set in a spreadsheet before you measure anything
One row per prompt, with columns for each assistant and each run. The set must not change between the baseline and the re-measure, because if you edit prompts mid-sprint you have no comparison at the end. Treat the spreadsheet as the instrument. Everything else in the sprint is read through it.
03
Run every prompt three times in fresh sessions and log answers verbatim
Use at least three different assistants, start a new chat for each run, turn off any memory or personalisation, and paste the full answer into the sheet rather than summarising it. Three runs matters because these systems are not deterministic, and a single run tells you about one sample rather than about your visibility.
04
Score each answer as named, mentioned, or absent
Named means you appear in the recommended set with a description. Mentioned means you appear somewhere in the answer but not as a recommendation. Absent means nothing. Count these across all runs to get a starting number. It will be low, and a low honest baseline is more useful than an impressive vague one.
05
Read the citations, not just the answers
Where an assistant links its sources, log every URL. The pattern across forty prompts tells you which specific pages the models are drawing on for your category, and those pages are usually comparison articles, documentation, and forum threads rather than vendor homepages. This list is the most valuable output of week two.
06
Find the gaps by comparing what is said to what is true
Sort the answers into three failure types. Prompts where a competitor is named and you are not, prompts where nobody credible is named at all, and prompts where you are described inaccurately. Each type needs different work, and the second type is the easiest win because the question is genuinely unanswered on the open web.
07
Publish pages that answer the prompt in the first two sentences
Week three. One page per gap, with the direct answer at the top before any context, a clear structure of question-shaped headings, and specifics rather than adjectives. Retrieval systems extract passages, so a page that buries its answer under four paragraphs of preamble is a page that will not be quoted from.
08
Get named on the third-party pages the assistants already cite
Your own site is only one input. Work the citation list from week two directly. Submit accurate entries to the comparison and directory pages that appeared, answer relevant forum threads honestly with your affiliation disclosed, and correct factual errors about you where the page allows corrections. This is slow and mostly manual, and it is the half most teams skip.
09
Re-run the identical prompt set in week four and record the delta
Same wording, same assistants, same three runs, fresh sessions. Compare named, mentioned, and absent counts against the baseline. Expect small movement or none at all in four weeks, since indexing and model updates run on timelines you do not control. The sprint produces the measurement system. The visibility follows later.
What this sprint actually produces
It produces a measurement system and a first round of content aimed at real gaps. It does not produce visibility in four weeks, and anyone selling that timeline is selling something else.
That is worth being blunt about, because the temptation with generative engine optimisation is to skip measurement and go straight to publishing. Teams that do this end up with twenty pages and no way of knowing whether any of them changed anything. The prompt set is the asset. The content is downstream of it.
Week one, building the prompt set
Write prompts the way buyers write them, which is in sentences, with constraints attached.
The four categories that matter are these. Category prompts, of the form what tools do X. Comparison prompts, naming a competitor. Problem prompts, describing a symptom rather than a category, which is where most early-stage buyers actually start. And constraint prompts, which add a budget, a team size, or a technical requirement, and which are the ones where a small specific product can beat a large general one.
Source the language from your sales calls. The gap between how founders describe their category and how buyers describe their problem is usually wide, and it is the single most common reason a product never gets named.
Week two, the audit
Run the set, three runs each, across at least three assistants, in fresh sessions with personalisation off. Paste answers in full.
Two numbers come out of this. Your share of answers, meaning the proportion of runs where you were named. And your citation map, meaning the set of URLs the assistants leaned on. Most teams find the second more surprising than the first. The pages that feed answers in a category are rarely vendor sites. They are comparison articles, documentation, community threads, and occasionally a single well-structured blog post from someone with no commercial interest at all.
| Failure type | What you saw | What fixes it |
|---|---|---|
| Competitor named, you absent | A rival appears consistently | Presence on the sources being cited |
| Nobody credible named | Vague or hedging answers | Publish the definitive page for that prompt |
| You appear but described wrongly | Stale pricing, wrong category | Find and correct the stale source |
Week three, filling the gaps
Write for extraction. The direct answer goes in the first two sentences of the page and of every section, because retrieval pulls passages rather than documents, and a passage that stands alone is a passage that can be quoted.
Beyond that, the rules are unglamorous. Use the buyer's words in the headings. Give numbers and specifics rather than adjectives. Date the page and keep it current. Say what your product is not good for, because comparative honesty is one of the few things that makes a page worth citing over the twenty vendor pages saying they are the best.
Then do the harder half, which is the third-party work. Correct entries, answer threads, get listed accurately. It is manual and there is no shortcut, which is precisely why it is worth doing.
Week four, and the honest expectation
Re-run the exact same prompts. Record the delta. Expect it to be small.
The reason to re-measure at four weeks anyway is that it establishes the cadence. Run the same set quarterly and the trend becomes readable within two or three cycles, which is the timescale on which this channel actually reports. What you are building is a repeatable instrument, and instruments are only useful when the readings are comparable.
The compounding argument is straightforward. The number of people who ask an assistant rather than a search box before they ever visit a vendor site is rising, and a product that is never named in those answers is invisible at exactly the moment a category is being shortlisted. That is a structural claim, not a statistic, and it is enough to justify a quarterly sprint.
Playbook questions
How is this different from SEO?
The assistant gives me a different answer every time. How can this be measured?
How long until anything actually changes?
Should I just get added to the best-of listicles?
What if an assistant says something factually wrong about us?
Do I need schema markup for this to work?
Or have us run it.
This playbook is yours to use. If you would rather not run it yourself, that is the entire point of Zway.
30 minutes. You leave with the plan either way.