Long sections on activity that produced nothing, described in the language of effort rather than outcome. And the most reliable indicator, a report you cannot disagree with, because it contains no specific claim to test.
It is also worth checking which assistant your customers actually use rather than assuming. The answer varies by profession, age and country far more than industry commentary suggests, and several businesses have built measurement programmes around a system their buyers never open. Adding one question to your enquiry form settles it in a fortnight and can redirect the whole effort.
The risk is scope drift into activity that is easy to report and hard to value. The protection is to have the retainer specify countable units: prompt set runs per month, listings audited, corrections submitted, pages published or rewritten, outreach attempts made.
Keep It Current and Say So Because retrieval happens at answer time, freshness carries real weight. A page updated this month can be cited this month, and a competitor can displace you simply by revising a page you have left alone for two years.
The honest framing first: nobody outside these organisations knows the selection logic, and the systems change without announcement. What follows is drawn from observable behaviour, visible citations and published research, which supports useful generalisations and does not support precision.
The Shared Architecture All three now commonly retrieve live sources rather than answering purely from training. Your question becomes one or more searches, a set of pages is fetched and read, and the answer is composed from what was read.
The Prompt Set, Unchanged The report opens with the prompt set used, versioned and dated, and a statement that it is identical to last month's. If it changed, the change is listed explicitly with a reason, and the previous series is kept alongside so comparisons remain honest.
The condition is that the output has to be yours to keep and act on elsewhere, including the prompt set. An audit that only makes sense inside that agency's retainer is a sales document with a price attached.
Good versions read like this: mention rate on evaluation prompts rose from two in fifteen to six in fifteen, which we attribute to the three directory corrections completed in week two, though a competitor also stopped publishing during the same period.
If the budget is substantial, add the earned coverage work, which is the slowest and most expensive component and the one you genuinely cannot do quickly on your own. Buying that first, before the cheap fixes are done, is the most common way money gets wasted in this field. generative engine optimization
A Claim About Causation, Stated as a Claim The most valuable paragraph in the report is the one that says what moved and why the agency believes their work caused it, phrased as a judgement rather than a fact.
In this case there is something real underneath. The plumbing of how people find suppliers has changed, and the work required has changed with it. Here is the whole idea explained without the acronyms, aimed at someone who wants to understand the decision rather than do the job. generative engine optimization
Keep a dated note of what you observed each quarter, including behaviour that later turned out to be temporary. The value is not in the individual observations, most of which expire, but in noticing how fast they expire. A team that has watched three of its confident conclusions become wrong within a year develops the right amount of scepticism about the fourth.
Observed behaviour leans toward breadth, pulling from a wider set of sources per answer than the others, and it cites forums, documentation and niche trade sources readily. It also appears comparatively responsive to freshness.
Pricing in this field is unusually opaque, partly because the work is new and partly because the absence of an independent scoreboard makes it hard for a buyer to tell whether they are getting value. That combination invites vague scoping.
Ask What They Cannot Measure A competent practitioner will volunteer limitations before you ask. Assistant answers vary between sessions. Referral attribution is inconsistent. Some assistants cannot be measured reliably at all. Sample sizes in the published research are small.
One overlooked cost is your own time. Every engagement in this field needs somebody inside the business to confirm figures, approve crawler changes and answer factual questions, and a plan that assumes this is free will stall. Budget a few hours a month explicitly and name the person, because the alternative is an agency waiting on answers and billing for a month in which little shipped.
Payment tied to a proprietary visibility score is worse, because the vendor controls the number and the methodology behind it. There is no independent scoreboard in this channel, which is precisely why performance pricing that works elsewhere does not work here.