How to move GEO testing beyond prompt tracking and prove real business impact

Prompt tracking alone cannot prove whether a GEO change is working. Here is how ecommerce teams can run controlled experiments across Google and AI surfaces.

Generative engine optimisation has created a fast-moving measurement market, but most teams are still trying to answer leadership questions with prompt tracking dashboards. Those dashboards can show whether a brand appeared in a sample of AI answers, yet they cannot connect a website change to a business outcome. Closed-loop testing on real page templates is the only way to prove whether a GEO change is worth rolling out.

The pressure is real. Leadership is asking whether the company is ready for AI discovery, whether the team is doing the right things, and how to move faster. Those questions deserve answers grounded in controlled experiments, not another checklist.

Why prompt tracking is not enough

Prompt tracking tools have proliferated across ChatGPT, Perplexity, Gemini, AI Overviews, and AI Mode. They sample prompts, log brand mentions, surface citations, and produce AI visibility scores. That information is useful for some jobs. It helps with debugging. It can show early warning signs when a brand disappears from common answers. It helps teams understand how a product is described in particular contexts.

What prompt tracking cannot do is prove that a website change made the business better. Prompts are personal, often long, and shaped by previous conversations. The same user can ask a follow-up that changes the whole context. Different users get different answers. The universe of possible prompts is effectively infinite. This is what practitioners are starting to call a search volume one world, where every query is a long-tail variant of itself.

The problem is not that prompt tracking is useless. The problem is that prompt tracking is being asked to do too much. It is not the same as measuring impact. It is not the same as proving that a change should ship across a large ecommerce site.

Why ecommerce is a practical testing ground for GEO

Ecommerce has a structural advantage in the AI discovery era. People still need a manufacturer or retailer to make and ship the product. An AI assistant can help with research, comparison, and shortlisting, but the transactional relationship stays with the retailer, marketplace, or brand. That means the risk is real, and the opportunity is real too.

More people will be buying more things online in the coming years, and many of those journeys will be shaped by some form of organic discovery. The interface may blend search engine and chatbot. The discovery process may be more conversational. But the commercial question is familiar: will the customer buy from you?

That makes product and category pages central. Large ecommerce sites run on scalable templates: product detail pages (PDPs), product listing pages (PLPs), category pages, internal search results, faceted navigation, buying guides, and related content blocks. Every one of those templates can be tested. A team can change something across a controlled group of pages, compare against a control, and measure what actually happened to both AI traffic and Google organic traffic. SearchPilot’s GEO A/B Testing platform is built around that same principle, and SearchPilot’s Merchant Center Testing work extends the same idea to product feeds and structured data.

How LLMs find fresh product information

AI discovery draws on two broad information sources. The first is training data, the information absorbed during the model’s training run. It shapes language, entities, relationships, and brand context but is largely fixed until the next training cycle. A team cannot walk into a leadership meeting with a strategy that amounts to waiting for the next model and hoping it likes the brand more.

The second source is retrieval. When a user asks a question, the model may pull in fresh information during the interaction, often described as retrieval-augmented generation, or RAG. In practice, the model may open dozens of background searches, read across the web, and synthesise an answer.

For ecommerce, retrieval is essential. A model cannot rely on old training data to answer questions about current stock, today’s price, active discounts, latest reviews, delivery options, product availability, new launches, or local availability. Product recommendations need freshness. That is where traditional search and AI discovery reconnect. AI systems need up-to-date information and reach for live pages, feeds, and search results. This is why product feeds, structured data, PDPs, pricing, availability, and product attributes all matter for how machines understand and recommend products.

What fan-out queries change about optimisation

When a user types one long prompt into an AI system, the model may break that task into many background searches. It may search for product comparisons, reviews, pricing, availability, best options for a use case, brand reputation, delivery details, and other supporting information. The user sees one answer. Behind that answer, many searches may have happened.

That changes how teams should think about optimisation. The old model was keyword-first, centred on which keyword to target, where the page ranks, and what the search result looks like. In AI discovery, the hidden fan-out queries are often where the real retrieval happens. Teams usually cannot see all of those queries. They may get clues, but they cannot treat the process as a clean list of keywords. That is one reason testing becomes more important. A team can make a change, then measure whether the change improved LLM referrals, Google organic traffic, or the net business outcome. Perfect visibility into every hidden query is not required to measure the effect of a change.

How to write a testable GEO hypothesis

A traditional SEO hypothesis usually works through one of three mechanisms: targeting new keywords, improving rankings for existing keywords, or changing the search result appearance so more searchers click through. GEO has analogues, but the language shifts.

A GEO hypothesis might aim to target new fan-out queries, improve visibility for existing fan-out queries, influence the summary an LLM returns, or make a page, product, or brand easier for the model to recommend. The fourth mechanism is interesting because it feels like conversion rate optimisation, except the converter is partly the machine. The question becomes whether the page has given the model enough information to confidently recommend the product. That can include product detail, comparison language, reviews, freshness, structured data, key features, FAQs, delivery information, stock information, and buying guidance.

A weak GEO hypothesis says, This might help AI visibility. A stronger GEO hypothesis says, Adding clearer product suitability information to PDPs may help models retrieve and recommend these products for more specific fan-out queries while also improving confidence in the AI-generated summary. That version gives a team something to test.

Where GEO testing actually happens on an ecommerce site

For large ecommerce sites, GEO testing happens on the same scalable surfaces that SEO testing already uses. That includes product detail pages, product listing pages, category templates, buying guide modules, comparison content, FAQs, review summaries, key feature summaries, internal linking modules, structured data freshness indicators, product feed-aligned content, and availability and delivery information. The mechanics look similar to SEO A/B testing. A team makes a change to a variant group of pages, compares performance to a control, and measures the result.

What changes is the journey being measured. In traditional search, a user might open several tabs, compare sources, read reviews, and check products before arriving at the site. Much of that research is visible across a set of searches and visits. In AI discovery, more of that research can happen inside the conversation. The model reads, compares, summarises, and narrows options before the user arrives. The site may only see the final click. That makes the click more valuable in some cases and harder to interpret.

Why GEO and SEO can disagree

Many GEO changes could plausibly help SEO too. More useful content, better structure, fresher product information, clearer summaries, stronger internal links, and better structured data all carry SEO hypotheses. That does not mean every GEO-positive change is SEO-positive. Most practical GEO work today still reaches AI systems through search-related retrieval, which creates overlap with SEO. Overlap is not the same as sameness.

A change can help an LLM understand and summarise a page while hurting Google organic performance. A change can make a page richer for AI retrieval while making it bloated, duplicative, or less effective in traditional search. This is where single-channel measurement becomes dangerous. A team could look only at LLM referrals, see a positive result, and roll the change out. If Google organic traffic falls by more in absolute terms, the business loses. That is the bigger risk with guessing: a visible win in one channel can hide a larger loss elsewhere. Measuring AI and Google channels together is the only way to see the net impact rather than celebrate one metric in isolation.

What the Omio test taught the industry

The clearest working example of this principle comes from SearchPilot’s work with Omio. Omio partnered with SearchPilot to share early GEO testing results with the wider industry. The key lesson was that GEO can be tested, and that GEO and SEO do not always move together. Two SearchPilot A/B tests on Omio illustrate the point.

In one test, adding brand USPs (unique selling points) increased LLM traffic by 18 percent. In another, adding structured key takeaways performed positively for LLM-driven traffic but was projected to reduce Google organic sessions by roughly 6.5 percent. Because the net business effect was negative, Omio chose not to roll the change out and developed follow-up iterations instead.

This is the practical value of testing. The AI result looked positive in isolation. The business result was not. Without measuring Google organic performance at the same time, the team could have shipped a net-negative change. It is also the strongest argument against treating GEO as a checklist. A tactic can be directionally plausible and still wrong for a specific site, page type, or business goal.

What prompt tracking can and cannot tell you

Almost every ecommerce team is now using one of the large prompt tracking or AI visibility tools. Almost every team also admits some version of not quite knowing whether to trust the data or what to do with it. That does not make the tools useless. The right comparison is rank tracking in traditional SEO.

Rank tracking is useful for debugging. It can diagnose problems, show movement, surface early warning signs, and help teams understand where visibility may be changing. But rank tracking is not the same as business impact. Prompt tracking has the same issue, with extra complications. Teams do not know the full set of prompts users are typing. Many prompts are unique. Answers are personalised. The model may use memory, previous conversations, location, and other context. A brand may appear for a prompt in one run and not another. A dashboard can only sample a small portion of reality. It can help teams form hypotheses and notice issues, but it should not be the main evidence that a GEO programme is working.

How to answer leadership questions about GEO

Executive attention on AI search has grown sharply. SEO leaders are fielding more questions from the C-suite than ever before, and the questions are sharp: are we ready, are we doing the right things, and how do we go faster. The honest answer is that nobody fully knows yet, but the teams that will move fastest are the ones that treat GEO as a testing programme, not a checklist.

That means framing each change as a hypothesis, picking a scalable surface to test it on, defining the control and variant groups, measuring both AI traffic and Google organic traffic, and deciding on rollout based on the net business result. When the data is clean, the conversation with leadership changes. Instead of arguing about whether a tactic is theoretically correct, the team can show what actually happened on the site.

FAQ

What is the difference between GEO and SEO?

SEO focuses on ranking in traditional search results. GEO focuses on influencing how large language models retrieve, summarise, and recommend content in AI answers. The two overlap because most AI systems still rely on search-related retrieval, but they can disagree on which specific changes help or hurt.

Why is prompt tracking not enough to measure GEO performance?

Prompt tracking samples a small portion of possible prompts, and prompts are personal, long, and shaped by previous conversations. It can show AI visibility movement, but it cannot connect a website change to a business outcome the way controlled A/B testing can.

What did the Omio GEO test reveal?

SearchPilot’s A/B tests with Omio found that adding brand USPs lifted LLM traffic by 18 percent, while adding structured key takeaways was projected to reduce Google organic sessions by roughly 6.5 percent. The result showed that GEO and SEO do not always move together, and that single-channel measurement can hide a net business loss.


This article summarizes reporting from searchpilot.com. See our editorial disclaimer for how our articles are produced.

🤖
Is your business visible to AI assistants?

Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.

Check Your Score →