Skip to content

Guide · Tracking

How to Measure the ROI of AI Search Optimization

Line illustration of a balance scale with a violet dollar coin on one pan and a starred AI answer card on the other, a robot at the pivot and a marketer weighing them beside a rising bar chart

AI search optimization ROI is the return on money spent getting a brand cited in AI answers, measured in citations and answer mentions gained rather than clicks won. That swap of units is the whole problem.

The surface you’re buying doesn’t reliably send traffic. So a report you build on sessions shows a number near zero. Your retainer gets cancelled on the strength of it.

Price the work in cost per citation, which is your program spend divided by net new citations in the period. Baseline four inputs first, on a prompt set with branded prompts stripped out. Then connect citations to pipeline through a four-link chain, and label each link by the confidence its data source supports. Two of those links you can evidence from first-party Google data, and the last two stay inference.

The click problem is measurable, which is why you should concede it early. When a Google AI summary appears, users click a standard search result in 8% of visits, against 15% with no summary present (Pew Research Center, July 2025, browsing data from 900 US adults, March 2025). Per the same data, clicks on a link inside the summary itself run at 1% of visits.

What ROI means for AI search optimization

GEO/AEO ROI is the return on AI-search work measured against a numerator of citations and answer mentions, not organic clicks.

Traditional SEO ROI has a clean chain: rank, click, session, conversion, revenue. Every link is instrumented, and the arithmetic is boring because the data cooperates.

AI search breaks the second link and keeps the rest. You can prove you were cited. You often can’t prove what the citation did.

So your unit of purchase changes. You are buying presence in an assembled answer rather than a position on a page, and presence is counted in citations and mentions.

That changes what your money should buy. In our analysis of 10,563 AI citations across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews, blog articles and listicles drove 60% to 72% of the total. Per the same dataset, vendor-owned corporate sites held 79% to 85% of the citation surface, against 0.2% to 2.6% for editorial publications.

So check one thing before you measure the return on a program: that it is buying the format which produces your numerator.

The four inputs you need before any math

Four numbers have to exist before ROI is calculable. They’re your inputs, not your outcomes, and none of them is the answer on its own.

InputWhat it countsWhere it comes fromWhat breaks it
Visibility scoreShare of tracked prompts where the brand appears at allPrompt-level tracking across enginesBranded prompts left in the set
Citation countTimes a brand’s own pages are cited as a sourceCitation logs per prompt, per engineCounting total instead of net new
Mention shareBrand’s share of all brands named in an answerBrand extraction across cited documentsUntagged competitors inflating the denominator
Average positionWhere the brand sits in a ranked answerPer-answer rank captureAveraging across unrelated prompt intents

The hygiene step comes before your baseline. Skipping it is the most common way one of these reports gets rejected. A prompt that names your brand returns near-total visibility by construction, because the model was handed the answer in the question. Pad your prompt set with branded prompts and the average climbs without anything real having moved.

Four numbered steps from prompt hygiene to baseline to a branded prompt, ending in inflated visibility near 100%

Stripping branded prompts is the default in our own tracking, before any metric gets read. Our own estimate, from demo and client calls across April to July 2026, is that roughly 90% of businesses come back at zero on the unbranded score.

But zero is a usable baseline. An inflated score is not.

In practice, that usually means:

  • Rebuild your set as unbranded commercial prompts.
  • Tag the real competitors so your mention denominator stays clean.
  • Group prompts by intent before you average position.
  • Record your starting values for all four before a single optimization ships.

Get the hygiene right before the tooling, then pick from a crowded category. Amadora AI, Semrush, Profound, Peec AI, Otterly.AI, Scrunch AI and AthenaHQ all track some version of these inputs. Amadora AI is our own product, and it tracks ChatGPT, Perplexity and Gemini daily on every plan.

The cost-per-citation formula

Cost per citation is total program spend divided by the net new citations that spend produced in the period.

Program spend ÷ net new citations = cost per citation

Program spend ÷ net new answer mentions = cost per mention

Run both. Citations track content and technical work. Mentions track third-party placement, and the two move on different schedules.

But the scoping matters more than the division. Net new means citations you did not hold at baseline. Citations you already had are not bought again this month. Counting them restates last quarter’s work as this month’s return.

Branded prompts come out, for the reason above. Two more exclusions keep your denominator honest:

  • Citations from pages that predate the program, unless the program is what got them cited.
  • Duplicate citations of the same URL inside one answer, which inflate volume without widening presence.

Program spend is the other half, and it’s the half agencies understate. It includes your tooling, analyst hours, content production, and any placement or PR cost.

Tooling is the only line with a published number. What prompt-level tracking costs is a fixed input you can quote exactly. That makes it the easiest place to start your sum.

Connecting AI visibility to traffic, leads and revenue

An attribution model for AI search is the chain connecting an answer mention to revenue, with each link labeled by the confidence its data source supports.

The chain has four links: answer mention → AI-surface impression → site visit → conversion. The first two are measurable. The last two are partly inferred. Publishing your chain with those labels attached is the difference between a report your client interrogates and one they accept.

What Search Console will and will not tell you

Google now reports on AI surfaces first-party, and what it reports sets the ceiling on your attribution model.

The generative AI performance report gives impressions for AI Overviews and AI Mode. Those impressions break out by page, country, device and date. It does not give clicks, click-through rate, or the query that triggered the answer (Google Search Console Help).

But the clicks aren’t missing from Search Console. They’re unlabeled. Traffic from AI features is folded into overall search traffic, under the Web search type in the Performance report (Google Search Central).

Read those two facts together and your ceiling becomes obvious. You get AI impressions with no clicks attached, and AI clicks with no AI label attached.

The two cannot be joined. So you cannot build a clean AI-versus-organic revenue split from Google’s own data. Anyone invoicing against one is selling an inference as a measurement.

What survives the audit is a four-tier label you apply to every number in your report:

  1. Measured, first-party. Citations, mentions, positions, AI impressions.
  2. Measured, indirect. Branded search volume and direct traffic, annotated against the dates work shipped.
  3. Self-reported. A how-did-you-hear-about-us field with an AI-assistant option, which is the only place a prospect tells you directly.
  4. Inferred. Everything else, including the revenue number your client most wants.

Report all four. Never promote a tier-4 number into a tier-1 claim because the slide looks better.

Worked example: an agency retainer ROI calculation

Four cost and revenue lines go into the calculation, and only one of them is a published fact:

  • Tooling. The real published price of the platform you use. Amadora AI is our own product, and its price sits in the table because we can state it exactly. Substitute whatever you pay.
  • Analyst time. Your loaded hourly cost times the hours the program consumes. The rate below is an assumption.
  • Net new citations. Your own measured number against the baseline. Never an estimate.
  • The GEO line you bill. Your rate.
LineOne clientTen clients on one plan
Tooling, Agency plan from $499/month (published price)$499.00$49.90
Analyst time, 10 hours at an assumed $120/hour$1,200.00$1,200.00
Program cost per client$1,699.00$1,249.90
Net new citations in the month (your number, 40 here)4040
Cost per citation$42.48$31.25
GEO line billed to the client (your rate, $2,000 here)$2,000.00$2,000.00
Gross margin per client$301.00$750.10

One line moved between those columns, and it wasn’t the labor. The Agency plan carries 300 prompts and multiple workspaces, so your fixed tooling cost amortizes across the client roster. Analyst time stays stubbornly linear.

Agency plan tooling cost spreading across three clients, set against analyst time that stays linear per client

Why this matters: at one client, a single scope creep erases that margin. The instinct is to cut the tooling line. That’s the wrong line to cut, because it’s the only one that gets cheaper per client as you add clients. Your labor line is what needs the productivity work.

Then there’s a second return line most agencies leave out. It sits on your own P&L rather than your client’s.

One agency reported closing about $35,000 in new monthly recurring revenue, using an AI visibility audit as its pitch asset. That was a mix of new clients and extended retainers (one self-reported account, May 2026, figure approximate, not a typical result).

But it names the right line. Work that wins and holds retainers returns to you directly, and a report measuring only your client’s pipeline undercounts its own value. The agency reporting and retainer workflows are where that motion gets documented.

What good ROI looks like, by team size

There is no single benchmark, because agencies and in-house teams have different return lines. Report one number for both and you misprice one of them.

The agency pass mark

Your pass mark is margin plus retention. That means cost per citation trending down month over month, with your GEO line surviving renewal.

Tooling amortized across enough clients to carry the fixed cost is what makes both possible. The retainer is your return. So measure whether AI search optimization ROI is holding your retainer, not whether it replaced your client’s paid search.

The in-house pass mark

You have no retainer, so your pass mark is share against a named competitor set, plus cost per citation against your existing cost per marketing-qualified lead.

If a citation costs you less than a lead and your mention trend is up, the program is working. That holds even while the revenue link stays inferred.

The target to set instead of a score

For the target itself, use mentions rather than a score.

In our tracking of the AI-visibility software category, roughly 100 to 250 mentions inside AI citations corresponded to a category top 10. Roughly 250 to 600 corresponded to a top three to five placing. (One category, measured May to June 2026. Thresholds moved over the period and are category-specific.)

You can baseline your unbranded prompt set and let it run a week before you read any trend.

What this framework will not tell you

Outside Google, nothing reports first-party. ChatGPT, Perplexity and Gemini publish no equivalent of an impressions report. So every number you get from those engines comes from sampling their public interfaces. Sampling carries a margin your report should state.

Net new citations move in steps, not curves. A single month can go flat while your underlying work lands. A monthly cost-per-citation figure read in isolation is noise rather than signal.

Instead of reading volume alone, check positioning when the two disagree. A brand with fewer total mentions can outrank one with more on a given intent. What third-party sources say the brand is decides which prompts it’s eligible for, so volume targets won’t fix a positioning problem.

Dmitry Chistov, co-founder and CEO of Amadora AI, argues that promising your client a target visibility score in a fixed window is indefensible. The result depends on niche, competition, and the foundation your client already has. His alternative: commit to two leading indicators instead. Own citation count moves with content and technical work, and brand mention count moves with placement.

And the last link stays inference. A client who audits your chain reaches tier 4, which is a conversation worth having in month one instead of month nine.

Common questions

How do you benchmark AI search optimization performance?

Benchmark competitively, not absolutely. Track your mention share against a tagged competitor set, on the same prompts and in the same window. An isolated visibility score means nothing without the set it was measured against. Re-read the comparison monthly and expect week-to-week movement.

How do agencies automate visibility reporting for clients?

Automate collection daily and reporting monthly. Prompt tracking, citation logging, brand extraction and per-answer rank capture run continuously. Your client-facing report is assembled on a fixed cadence, with work-shipped dates annotated against the metric movement. Filter branded prompts out of any shared view first.

How does AI search ROI compare with traditional SEO ROI?

The payback period is longer and the attribution is weaker. SEO ROI resolves to instrumented clicks, while AI search ROI resolves to citations plus an inferred revenue link. Report them on separate units rather than blending them into one ROI number, because the blend hides which half of your figure is a measurement and which half is a judgment.