Skip to content

Day 106: Do Not Let the GEO Brief Mark Its Own Homework

A GEO programme can improve exactly where it was told to improve and still leave the buyer problem largely untouched.

The team begins with a fixed set of commercial buyer questions. The questions are sensible. They expose a real gap: the company is absent, miscategorised, routed towards the wrong provider type, or described with a weak next step. The team uses those questions to brief the work. Pages are clarified. Comparison language is sharpened. Offer boundaries are made easier to understand. The same questions are checked again.

The result looks better.

That is useful like-for-like tracking. It is not independent evaluation.

For CMOs, Marketing Directors, and founders, the risk is simple: do not approve a content sprint, retainer, or success claim that only improved on the buyer questions used to shape the work. Before remediation starts, reserve a small family of unseen but commercially equivalent questions. Use the original set to guide the intervention. Use the holdout family to test whether the improvement travels to adjacent wording, constraints, category cues, and buying routes.

The same instrument should not be the brief and the proof

A fixed question set is valuable because it gives the team a stable instrument. It makes the problem concrete. It lets the team compare like with like. It prevents vague complaints about AI visibility from turning into a sprawling content backlog.

But the instrument has a second effect: it teaches the team what to optimise for.

If a buyer question says, "Which firms help B2B SaaS CMOs understand whether AI recommendations are sending prospects towards the wrong kind of provider?", the public changes will naturally start to answer that exact situation. The page title may echo the role. The comparison section may name the wrong-provider risk. The offer copy may describe the diagnostic in that language. The sales note may frame the next step around that question.

That can be the right work. The problem appears when the same question then becomes the success evidence. The team has not only measured the market. It has also supplied the vocabulary, constraint, buyer role, and buying route that guided the fix. A better result on that original question may show that the intervention matched the brief. It does not necessarily show that the company has become more legible when a buyer asks the same commercial question differently.

In other words, the brief may be marking its own homework.

A compact mini teardown

Imagine a specialist firm wants to understand whether answer-led research can identify its diagnostic offer for a CMO worried about low-fit demand entering sales.

The intervention question is precise:

"Who can help a B2B SaaS CMO diagnose whether AI recommendations are sending prospects towards the wrong type of provider?"

The initial observation is weak. The answer talks about monitoring tools, broad SEO agencies, and internal analytics. The firm is not clearly described as a diagnostic partner. That finding supports practical public work: clarify the offer page, explain when a diagnostic is different from monitoring software, name the sales-conversation consequence, and add a next step.

After the work, the same question improves. The firm is mentioned more clearly, or the category route is described more accurately, or the answer explains the diagnostic/service distinction better.

Good. The intervention has moved the original observation.

Now ask the holdout family that was reserved before the rewrite:

  • "Which partner can help a marketing director understand why answer engines are creating poor-fit enquiries before sales calls?"
  • "How should a B2B founder check whether AI answer surfaces are steering buyers towards software when the real need is advisory diagnosis?"
  • "What kind of specialist should we speak to if prospects arrive from AI research with the wrong expectation about our service?"

Those questions keep the commercial situation comparable: same buyer problem, same broad market, same buying stage, same intent to find a credible route. But they vary the role, wording, category label, constraint, and route framing.

If the improvement travels, the public material may have become more broadly legible. If it does not, the team may have tuned the visible answer to one phrasing while leaving adjacent buyer language unsupported.

That is the useful distinction.

How to build a small holdout family

A holdout check does not need to become a large research programme. It needs to be designed before the intervention, protected from the copy brief, and close enough to the buyer reality to be commercially meaningful.

Start with the buyer situation, not the wording. Name the role, problem, market, buying stage, and decision the question represents.

Then create two sets.

The intervention set contains the questions the team will use to diagnose the gap and guide public changes. These can be inspected closely. They can shape page priorities, comparison language, offer boundaries, evidence notes, and sales enablement.

The holdout buyer-question family is reserved. It should not be used as source language for the rewrite. It should express the same commercial buying moment while changing several surface features:

Keep comparable Vary deliberately
Buyer problem Wording and synonyms
Buyer role CMO, Marketing Director, founder, sales leader, or procurement view where relevant
Market and company type Category label or provider type
Buying stage Constraint, risk, or urgency
Commercial intent Route framing: specialist, tool, agency, internal team, or no-action check

The goal is not to hide a trick question from the team. The goal is to stop evaluation becoming a mirror of the remediation brief.

What the holdout can and cannot prove

The holdout family should be treated with restraint.

It can show whether a public change appears to travel beyond the exact questions that shaped it. It can expose over-specific copy, reveal dependence on the supplier's preferred vocabulary, and help a CMO decide whether a reported improvement is strong enough to justify more work.

It cannot prove real buyer behaviour. It cannot prove demand. It cannot prove attribution, pipeline movement, market share, or causal platform response. It does not show that buyers asked those exact questions, saw the same answers, trusted them, or acted on them. It is a bounded observation design.

That boundary makes the method more useful, not less.

A buyer does not need a supplier to pretend that a small holdout is a controlled market experiment. They need to know whether the success claim has escaped the closed loop. If the only positive movement appears on the questions used to brief the work, procurement should ask for more caution. If adjacent commercially equivalent questions also improve under recorded conditions, the case for continued remediation is stronger, while still limited.

The useful claim is modest: under recorded conditions, the public changes improved the original intervention questions and also appeared to travel to reserved adjacent buyer-question variants.

Record the context before comparing

A holdout method fails if the observation record is sloppy.

For each intervention and holdout question, record the surface, date, access context, geography or market setting where relevant, visible sources where available, and any obvious source-state changes. Treat ChatGPT, Claude, Perplexity, Gemini, Google AI features, search results, review sites, directories, and other public contexts as different observation surfaces, not one interchangeable channel.

Repeat responsibly. One before-and-after answer can start the discussion, but it should not carry the whole management claim.

Google needs the familiar caveat as well. Google's AI features rely on core Search ranking and quality systems. A weak holdout result there should not be blamed on missing llms.txt, special AI markup, arbitrary chunking, or over-focused structured data. Inspect whether useful, relevant, accessible, high-quality public material exists for the buyer problem and route.

The context record protects both sides of the engagement. The supplier cannot overclaim a convenient re-run. The buyer can see which observations were intervention checks, which were reserved holdouts, and which claims remain outside the evidence.

The procurement implication

A credible GEO proposal should be willing to name its acceptance test before the work starts.

Not a universal score. Not a guaranteed uplift. A simple separation:

  • these buyer questions will diagnose and guide the intervention;
  • these commercially equivalent holdout families will be reserved;
  • these surfaces and access contexts will be recorded;
  • these limits will remain attached to the results;
  • these claims will not be made from the evidence.

That structure helps a CMO buy the work with less theatre. It reduces the chance that a supplier optimises for the exact report questions, returns with a better screenshot, and calls it market progress.

The practical close is not complicated.

Before the rewrite, reserve the unseen questions. After the rewrite, check both sets. If only the intervention set moves, treat the result as a brief-fit improvement. If the holdout family also moves under recorded conditions, treat it as stronger but still bounded evidence that the public explanation has become more legible.

Do not let the question set that shaped the work become the only judge of whether the work succeeded.