Ask what was done, not what happened. If listings were corrected, pages rewritten and outreach attempted, and the numbers are still flat, that is information about the market. If none of it happened, the numbers were never going to move.
One check is worth running independently once a quarter, without telling anyone. Take ten prompts from the agreed set, run them yourself in a signed out session, and compare what you find against the most recent report. Broad agreement is reassuring. A consistent gap in the agency's favour is the single most informative finding available to you, and it is not something a report will ever surface.
State What You Sell in Concrete Terms Price range, lead time, geography, capacity, what you decline. This feels commercially sensitive and it is the material that makes your pages quotable, so an agency that does not have it will write vague content by necessity.
If you will not name competitors, say that too, and understand it removes the highest performing content format from the plan. Better to have that argument in the brief than to have a comparison page written and then killed.
The Signals That Mean Something Four things are hard to fake and worth watching closely. Your own pages beginning to appear in cited sources, which is directly observable in any assistant that shows citations.
Performance and Score Based Models Both sound aligned and both create problems. Payment tied to mentions creates pressure to shape the prompt set toward questions you already win, which is measurable improvement that means nothing.
Screenshots of favourable answers with no run count, which say nothing about how many attempts produced them. Impressions or traffic from unrelated channels included to fill a report. And activity described in the language of effort, such as ongoing optimisation, with no countable output attached.
Distinguish between a supplier who is failing and one who is reporting badly, because the remedies differ entirely. Ask for the raw answers and read them yourself before deciding. It is not unusual to find that sound work has been buried under a dashboard nobody understands, and fixing the reporting is far cheaper and less disruptive than replacing a team that is actually doing the job.
What to Spend Where If the budget is small, buy the audit and do the listings work yourself. Correcting your presence on the sources that already get recommended by ai cited is the highest return activity available and it requires attention rather than expertise.
What Has Not Changed It is worth being clear about the continuities, because the change is regularly oversold. Organic search still delivers the larger share of traffic for most businesses. Crawlable, fast, well structured sites still win. Content that genuinely answers a question still outperforms content that does not.
Be wary of proposals where the largest line is content production. It is the easiest work to scale, the easiest to bill and the least likely to be the constraint, particularly before a baseline exists. A proposal weighted toward diagnosis, technical fixes and third party corrections is usually cheaper and almost always sequenced better.
Run each prompt at least three times. Assistants vary their answers between runs, and a single result is a sample rather than a finding. Record the full text of each answer and every source cited, not a summary.
What Is Likely Next Forecasting specifics here is a good way to be wrong in public, so two general observations will do. First, the direction of travel has been consistent for a decade: interfaces keep absorbing more of the work the user used to do, and each absorption removes a category of click.
Then load your key pages with scripts disabled. Whatever remains is roughly what a retrieval system sees. If your product specifications, pricing or service areas vanish, that content needs to exist in the server rendered HTML.
Because there is no independent scoreboard in this channel, an engagement can run for a year on the strength of a number the supplier produces. That is an unusual amount of trust to extend, and it makes knowing what to check more important here than in any other marketing channel.
Making Any Model Safe Four clauses do most of the protective work regardless of structure. The prompt set and baseline archive belong to you and leave with you. Raw answers ship with every report. Scope is stated in countable units. And there is a defined review point with agreed criteria before the contract auto renews.
The guard against this is boring and effective. Change one substantial thing at a time where you can, record what you did and when, and note the alternative explanations alongside your conclusion. Attribution in this channel is genuinely hard, and a team that admits that will make better decisions than one that produces a confident causal story after every movement.
A reasonable formulation: after two quarters, we expect movement in mention rate on buying intent prompts, improvement in the accuracy of how we are described, and new citations from the sources our baseline showed matter. If none of those move, we will treat the approach as unsuccessful.