Generative Engine Optimization Agency selection should start with a diagnosis, not a sales call.
Your brand can disappear from AI-generated answers for four different reasons: retrieval, entity recognition, corroboration, or measurement. Each problem requires a different solution. Fixing the wrong one wastes time, inflates costs, and leaves your visibility unchanged.
Most companies hire an agency before they identify the real constraint. That gives the agency control of the diagnosis. Content agencies recommend more content. Technical agencies focus on crawling and rendering. PR firms push outreach. Each firm sells the service it already knows how to deliver, even when your business needs something else.
This guide puts you back in control.
In two hours, you can identify the issue that limits your AI visibility, determine which expertise you actually need, and judge every proposal against a clear specification. You will build a shorter shortlist, reduce the cost of your first engagement, and avoid paying for a year of work that never addresses the real problem.
You will not find an agency ranking here. Most published lists reward sponsorships, affiliate relationships, or paid placement. You need a diagnosis, not another list of vendors.
The four reasons you are missing from AI answers
Every credible GEO engagement traces back to one of these. Diagnosing which is yours is most of the work.
1. Retrieval. They cannot fetch or parse you. Your content requires JavaScript to appear, your CDN treats retrieval bots as a threat, or a bot-mitigation rule someone switched on in 2023 is quietly refusing the fetchers that decide citation eligibility. This one is cheap and fast to fix, and it is more common than agencies admit.
2. Entity. They cannot resolve who you are. Your company name, category, founding year, executive names, and product descriptions disagree across your site, LinkedIn, Crunchbase, G2, and Wikidata. Faced with contradiction, a model hedges or substitutes a competitor it understands better.
3. Corroboration. Nobody else vouches for you. Ask any assistant which vendor is best in your category and watch what gets cited. Review platforms, trade roundups, comparison pages, forum threads. Your own site appears as a supporting source at most. If you are absent from the 15 to 25 domains that keep surfacing, on-site work has a ceiling you will hit in month three.
4. Measurement. You cannot tell which of the above is true. Without a fixed prompt set and a baseline, every subsequent claim about progress is unfalsifiable. Plenty of retainers are sold specifically into this gap.
Retrieval and entity problems are engineering and reconciliation work. Corroboration is earned media. They are different skills, staffed by different people, priced differently. Buying the wrong one is the single most expensive mistake in this category.
Run the diagnosis yourself first
Two hours, no budget, no NDA. Do this before you take a single sales call.
Step 1: Find your real competitive set
Write down ten prompts your buyers would actually type. Not keywords. Full questions with constraints, the way people talk to assistants: “best payroll software for a 40-person construction company that needs multi-state tax filing.”
Run all ten through ChatGPT, Perplexity, and Google AI Mode. Log two things per prompt: which brands get named, and every domain cited.
That domain list is the finding. It is your genuine competitive set, and it is almost never the set of companies you thought you were competing with. Most of it will be third-party publishers you have no relationship with. A June 2026 analysis of 131,514 URL-grounded citations across 128 brands found that 85.7 percent pointed to third-party sites, compared with just 14.3 percent to brand-owned domains.
Step 2: Check crawler access, and know the three categories
Here is where most audits go wrong, including paid ones. AI crawlers are not one population. They split into three groups that answer three separate business questions:
- Training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended). Blocking these is a licensing and IP decision. It has essentially no effect on whether you get cited in answers today.
- Search and retrieval crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot). These determine citation eligibility. Blocking them is how brands make themselves invisible without meaning to.
- User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User). A human is waiting on that request right now.
A single robots.txt rule aimed at “AI bots” collapses three separate business decisions into one answer, and that answer is usually wrong. Treat robots.txt as a declaration of preference, not proof of enforcement: Cloudflare notes that compliance is voluntary, while its AI Crawl Control and WAF rules can technically block requests. Check all three layers separately.
Cloudflare’s behavior-based controls for Search, Agent, and Training traffic are already available, with new defaults taking effect on 15 September 2026. For newly onboarded domains, Training and Agent bots will be blocked by default on pages displaying ads, while Search bots will remain allowed. On the same date, Cloudflare will begin applying the most restrictive selected policy to mixed-purpose crawlers.
That means sites configured to block Training, including those still using the legacy “Block AI bots” setting, may also block dual-purpose crawlers such as Googlebot, Applebot, and BingBot unless the owner opts out or changes the policy before rollout. Existing WAF rules can independently override an “Allow” choice, so audit Security Settings and custom WAF rules rather than relying on robots.txt alone.
Step 3: Diff your own entity
Open five properties side by side: your homepage, your LinkedIn company page, Crunchbase, your primary review platform profile, and your most recent press release. Read the one-sentence company description on each.
If they disagree on what category you are in, who you serve, or how big you are, you have an entity problem. It is tedious to fix and frequently the highest-leverage item on the list.
Step 4: Pull your free Google-side baseline
Google now publishes a Search Generative AI performance report inside Search Console, launched on 3 June 2026, showing impressions for AI Overviews and AI Mode as a view separate from classic search results. It reports impressions, pages, countries, devices, and dates. It does not report clicks, and it is rolling out to properties gradually.
This matters for your budget. Part of what agencies were charging to estimate is now available to you directly for Google’s surfaces. It tells you nothing about ChatGPT, Perplexity, or Claude, which is exactly where third-party tracking still earns its cost. Know which half you are paying for.
Match the bottleneck to the kind of firm
| Your constraint | What you are actually buying | Who is built for it |
|---|---|---|
| Retrieval | Engineering hours and log analysis | Technical SEO specialists, in-house dev time |
| Entity | Reconciliation and structured data | A consultant, or two weeks of your own team |
| Corroboration | Relationships and outreach capacity | A digital PR-led generative engine optimization agency |
| Measurement | Prompt-set design and tooling discipline | Independent generative engine optimization consultants |
Read that table honestly and you will notice something. Three of the four rows can be handled without an agency retainer if you have a competent technical SEO and a writer with real subject-matter access.
The one you cannot build quickly is corroboration. Earning a place in someone else’s roundup runs on their editorial calendar, not yours, and it depends on relationships that take years to accumulate. That is the capability worth agency rates. Everything else is negotiable.
GEO, AEO, LLMO, SXO: the label tells you nothing
You will see the same service sold as generative engine optimization services, answer engine optimization, LLM optimization, AI search optimization, and relevance engineering. Some firms describe themselves as generative AI search engine optimization agencies, others as generative search engine optimization companies. The vocabulary has not settled and it is not a quality signal.
Google’s own AI-optimization guidance is blunt about this: from its perspective, optimizing for generative search is still SEO, with no separate technical requirements beyond normal Search eligibility. It also tells site owners to skip the hacks, naming unnecessary AI text files such as llms.txt directly, and confirms that Google Search ignores them. Other systems may choose to read such a file, so treat it as a cheap experiment, never as a strategy.
The useful framing: roughly two thirds of generative engine optimization services are strong SEO fundamentals applied with retrieval in mind. The remaining third is new work, namely prompt-set construction, citation-share measurement, source-level competitive analysis, and off-site corroboration aimed at the specific domains models pull from. Judge candidates on that third. The rest they should already be able to do without a premium.
What belongs in the scope, and what is filler
Worth paying for:
- A locked prompt set of 100 to 300 buying-intent questions, built from sales transcripts and closed-won deals rather than a keyword tool
- A citation-share baseline for you and three to five named competitors, produced before you sign
- Source-level attribution showing which third-party domains earn your citations
- A technical retrievability fix list with named owners and dates
- Earned placements on the domains your Step 1 audit actually surfaced
- Entity reconciliation across properties you control and those you can influence
Usually filler:
- An llms.txt file presented as a centrepiece. It costs an hour. It is not a strategy.
- Bulk FAQ schema on pages nobody retrieves
- “AI engine submission,” which is not a mechanism that exists
- A proprietary visibility score with no source-level detail underneath it
- Content volume targets. Twenty forgettable posts lose to three that other people cite.
Nine questions that separate practitioners from repackagers
Send these before the call. Watch where the hesitations land.
- How did you build my prompt set? Good: sales transcripts, support tickets, closed-won deals. Bad: a keyword tool with “best” prepended.
- What is my citation share right now, and against whom? Anyone serious can produce a baseline before you sign. If they cannot, they have not looked.
- Which surfaces do you measure, and how often? ChatGPT, Google AI Overviews and AI Mode, Perplexity, Gemini, Copilot, and Claude behave differently. One aggregate “AI visibility” number hides everything useful.
- How do you handle answer variance? The same prompt returns different answers to different users on the same day. If they are not sampling repeatedly, their trend line is noise with a smooth curve drawn through it.
- Which third-party domains dominate my prompt set, and what is the plan for each? This is the question that sorts strategists from resellers. You want named domains and named tactics.
- What percentage of this engagement happens off my website? Under 30 percent means you have bought an SEO retainer with a new cover.
- Show me a redacted month-four report from a current client. Month one proves nothing. Month four shows whether anything moved and whether they reported it honestly when it did not.
- What did you try in the last six months that failed? Real practitioners have a list. Repackagers have a case study.
- What do I own at exit? The prompt set, the raw data, the schema, the content, the media relationships. Get it in writing before the first invoice.
What generative engine optimization services cost
Every pricing benchmark currently in circulation is published by a company selling GEO, and the ranges span from a few hundred dollars a month to six figures. That spread is not a market rate. It is an admission that the market has not settled. Treat published figures as directional at best.
The closest external check is Clutch’s July 2026 GEO agency directory. Listed providers report minimum project sizes ranging from $1,000-plus to $50,000-plus, with hourly rates ranging from under $25 to more than $300. But these are agency-profile figures, not audited prices from comparable GEO engagements, and Clutch discloses that it may earn fees from some placements. Treat the range as a proposal sanity check, not an independent market rate.
The structural patterns are more reliable than the numbers.
Audit and baseline projects are one-time engagements producing a prompt set, a citation baseline, a technical fix list, and a source-target map. This is the correct first purchase for most companies. Insist that the raw prompt set and data are yours.
Execution retainers cover monthly measurement, technical work, content, and some earned media. Most generative engine optimization companies sell here. The variable that matters is how much earned media is included, because that is the expensive part and the part most likely to be quietly missing.
Hybrid and performance-linked models tie part of the fee to citation-share movement. Workable, provided the baseline is documented before work begins and the prompt set cannot be swapped for easier questions later.
Three sizing tests. Divide the monthly fee by the number of earned placements they will commit to in writing; if that number looks like agency margin rather than media work, you are paying for reporting. Ask whether tracking-tool licences are bundled or passed through. And budget against your deal size, not your traffic, because a company closing $60,000 contracts needs a handful of citations to justify a large retainer while a low-ticket brand needs volume.
Agency, consultant, or in-house?
Hire an agency when you need earned media at volume and lack the relationships to get it. That is the capability you cannot build in a quarter, and it is the main defensible reason to pay agency rates.
Hire a consultant when you have a competent content and dev team but nobody who knows what to point them at. A senior generative engine optimization consultant on a modest retainer can build the prompt set, design the measurement, and direct your existing people. Frequently the best value available, and the option agencies rarely raise.
Keep it in-house when you have a technical SEO you trust with schema and crawler configuration plus a writer with real access to subject-matter experts. Buy the tracking tool, spend two weeks on the baseline, and you have covered most of the retrieval and entity work. What you will still lack is reach.
Most mid-market teams land on a hybrid: outside help for measurement design and off-site work, internal people for anything requiring product knowledge. That structure also survives a budget cut better than one large retainer.
Structure the first engagement as a 90-day pilot
Send the identical brief to three finalists and compare what comes back.
Days 1 to 30. Build and freeze the prompt set. Establish the baseline across every surface your buyers use. Ship the technical fix list. Deliverable: a ranked list of target domains with a named plan for each.
Days 31 to 60. Execute the top five technical fixes. Publish two to three assets built to be cited, meaning original data, named experts, and specific numbers rather than adjectives. Pitch the first wave of roundups and comparison pages. Deliverable: first measurement re-run against baseline.
Days 61 to 90. Second measurement pass. Review source-level attribution. Deliverable: an honest read on which lever moved, which did not, and what they would do differently with full budget.
Judge the final readout on candour as much as results. An agency that reports flat citation share while explaining which leading indicators did move is worth more than one that finds a way to declare victory. If a firm resists a 90-day pilot with hard deliverables, that refusal is your answer.
Realistic timelines: retrieval fixes can shift results in two to six weeks. Corroboration takes a quarter or more.
Realistic timelines: retrieval fixes can become crawlable and indexable within days to weeks, Google says recrawling alone may take anywhere from a few days to a few weeks, but that does not guarantee immediate inclusion or citation. For corroboration, treat one quarter as a minimum planning window, not an established benchmark. A July 2026 review of 45 GEO studies found no stable, longitudinal, cross-platform causal evidence establishing how quickly organic citation visibility changes.
Red flags worth walking away from
- Guaranteed citations or guaranteed placement in ChatGPT. Nobody controls model output.
- Work starting before a baseline exists.
- Mass listicle placements on domains you have never heard of. Same playbook as link farms, same ending.
- Reporting that is one aggregate score with no source breakdown.
- Astroturfing communities. It gets detected, bans are hard to undo, and screenshots last forever.
- A pitch built around traffic recovery. Assistants send fewer clicks by design. Your goal is influence over the answer, which shows up as branded demand and better-informed sales calls, not restored sessions.
The takeaway
Diagnosis before shopping. Every hour you spend identifying your own bottleneck removes a month of paying someone else to guess at it.
Do this today: pick five high-intent buying prompts, run them through ChatGPT, Perplexity, and Google AI Mode, and write down every domain cited. Then check whether your retrieval bots are blocked at the CDN, and pull your Search Console generative AI report if your property has it.
By tomorrow you will know which of the four constraints is yours. Take that to the first sales call and the conversation changes shape entirely. You stop being pitched and start buying against a specification.

