Ahrefs found in July 2025, across 15,000 long-tail prompts and four assistants, that around 80 percent of cited pages did not rank for the original query at all. If citation and ranking were the same thing, that number would be close to zero. generative engine optimization
Weeks Three and Four: The Access Findings A technical report covering crawler permissions, what the relevant agents actually receive from your server, whether bot management is interfering, and what survives on your key pages with JavaScript disabled.
Pricing in this field is unusually opaque, partly because the work is new and partly because the absence of an independent scoreboard makes it hard for a buyer to tell whether they are getting value. That combination invites vague scoping.
What Is Likely Next Forecasting specifics here is a good way to be wrong in public, so two general observations will do. First, the direction of travel has been consistent for a decade: interfaces keep absorbing more of the work the user used to do, and each absorption removes a category of click.
What Should Not Have Happened Yet A large volume of new content. Twenty published articles by month three usually means the baseline was not used to direct the work, and the pages were commissioned before anyone knew which questions mattered.
Testing too rarely means you find out about a problem a quarter after it started. Testing too often means drowning in variance that looks like signal and reacting to noise. Both failures are common and the second is more expensive, because it produces work.
Freshness Counts More Than You Expect Because retrieval happens at answer time, a page published or updated this week can be cited this week. This is a meaningful difference from ranking systems where authority accrues slowly.
Alongside it, the first rewritten pages. Not a volume of new content, but your most commercially important existing pages restructured to answer directly and to carry specifics. You should be asked to confirm figures, since nobody outside your business can verify a lead time or a price range.
If you must change the prompt set, add new prompts as a separate cohort and keep the original series running unchanged. Editing the instrument retrospectively destroys the comparison you have been building.
This means a single answer is a sample. Being absent once is not evidence of a problem and being named once is not evidence of success, and treating either as a result is the most common analytical error in this field.
A useful way to think about the sequence is that each stage moved a task from the user to the interface. First the fact, then the summary, and now the comparison. Each move removed a reason to visit a website, and each was followed by an industry insisting the change had been overstated. It is reasonable to expect the pattern to continue rather than to stop at a convenient point.
For roughly twenty years the arrangement was stable enough that an entire industry could be built on it. You typed a query, you got a ranked list, you formed your own opinion by comparing a few of the results, and businesses competed for position in that list.
This is also why review volume and recency show up so consistently in what gets cited. A platform with forty recent accounts of working with you is more informative than your own page saying customers love you, and it is treated accordingly.
Also watch what happens to your citations over time rather than checking once. A page that earns a citation and then loses it usually has a fresher competitor rather than a technical problem, and the fix is updating your figures rather than rewriting the page. Because retrieval runs live, that maintenance is cheap and it is the difference between a page that keeps earning and one that quietly stops.
Second, the businesses that have weathered each stage best are the ones that were not dependent on a single channel. That was true when featured snippets arrived, it was true through every core update since, and it is true now.
When to Test More Often Three situations justify a tighter loop. During an active campaign where you need to attribute a specific change, weekly runs on a subset of prompts are reasonable, provided you accept the variance.
Payment tied to a proprietary visibility score is worse, because the vendor controls the number and the methodology behind it. There is no independent scoreboard in this channel, which is precisely why performance pricing that works elsewhere does not work here.
One thing to establish in week one is where everything lives. The prompt set, the baseline archive, the raw answers and the correction log should sit somewhere you control from the beginning rather than in the agency's systems. Retrieving them later is a negotiation. Having them from the start is an administrative decision nobody objects to at the outset.
How to Run the Ninety Day Review Ask three questions. Can you show me the prompt set is unchanged. Can you show me the raw answers. What specifically did you do, and which of the changes do you believe caused which movement.