AI & Brand Strategy

Does generative engine optimisation work? The 2026 evidence is thinner than the pitch

By Vantage Branding·Reviewed by ·15 September 2026·14 min read

Generative engine optimisation is the practice of changing content so that AI answer engines such as ChatGPT, Perplexity, Gemini and Google AI Overviews are more likely to retrieve and cite it. Does it work? Partly, and much less than it is sold as doing. The first systematic review of the field, published in July 2026 and covering 45 studies, found that no reviewed technique shows a stable, cross-platform causal effect on discoverability.

In July 2026 the practice acquired something it had been sold without: a systematic review of its own evidence base. The review covered 45 studies and concluded that no technique it examined shows a stable, longitudinal, cross-platform causal effect. That finding does not make the category worthless. It makes most of what is currently sold under its name unsupported, and it points clearly at the few things that hold.

What generative engine optimisation claims to do

Generative engine optimisation is a set of content and technical practices intended to raise the probability that an AI answer engine cites a given page. In its strong form, vendors claim it works like search engine optimisation did in 2010: apply the recipe, gain the visibility. That framing is the source of most of the confusion, because the two systems fail in different places.

An answer engine does two separable things. It retrieves a candidate set of documents, then it composes an answer from that set and decides which documents to credit. Almost all published GEO research operates on the second step, with the candidate set already fixed. The techniques are then reported as though they raised visibility, when what they raised was the odds of being quoted from a shortlist the technique had no part in reaching.

For the question of whether a brand should fund GEO or traditional search optimisation, and how the two disciplines divide, see our analysis of GEO versus SEO and what the data says a brand should fund now. This article asks a narrower and more awkward question: of the techniques sold under the GEO banner, which ones survive testing?

The 40% figure everyone quotes comes from one experiment

Almost every GEO pitch deck in circulation carries a version of the same number: optimisation lifts visibility by around 40%. The July 2026 survey traced it. The figure comes from a single 2024 benchmark configuration, in which one visibility metric rose from 19.3 to 27.2 when quotations were added to a document. That is a relative gain of roughly 41% on one metric, in one experimental setup, for one technique.

The survey's evidence hierarchy lists the claim that GEO increases visibility by 40% as rejected as a general claim, with the rationale that the figure is a relative maximum on one metric under a specific configuration. It is not a finding about brands, categories or platforms. It is a finding about what happens to a scoring function when a quotation is pasted into a document that a system has already decided to look at.

This is the pattern that repeats throughout the literature. The mechanism is real. The generalisation is not. Clinical research has a word for this gap: a compound that works in vitro has not been shown to work in patients, and the history of medicine is largely the history of that distinction being learned expensively. GEO's foundational results are in vitro. The trials came later, and they came back weaker.

At this stage, claims about GEO return on investment clearly outstrip the academic evidence.

That sentence is the survey's own, and it is the most honest line published about the category this year.

Forty-five studies found no technique with a stable effect

The survey's central conclusion is worth quoting in full, because it is routinely softened in summary: already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behaviour.

Read that carefully. Already-retrieved content. The evidence supports influencing what happens to a document after a system has found it. It does not support the claim that these techniques get a document found.

The survey grades individual techniques by the strength of the evidence behind them. Exhibit 1 reproduces those gradings and the conditions the paper attaches to each. The conditions are not decoration. In several rows they reverse the practical advice.

Exhibit 1: The GEO evidence grade. Techniques commonly sold as generative engine optimisation, graded by the strength of published evidence. Gradings and cautions follow the July 2026 arXiv survey of 45 studies.
TechniqueEvidence gradeCondition or caution
Query–document relevanceStrong in controlled settingsPrimary determinant of the first citation; addresses genuine information needs
Position in the context windowStrongConditional on the document already being retrieved
Extractable evidence (quotations, statistics)Moderate to strongVerifiable figures, definitions and comparisons; numbers are the unit of citation
Recency, prices and datesModerateUseful for time-sensitive or commercial queries, not universal
Document structureModerate and heterogeneousTest headings, tables and fields without assuming the direction of effect
Fluency and simplificationWeak to moderateDomain and engine specific; optimise for the user first
Authoritative toneWeak and unstableMay conflict with credibility; do not conflate confidence with evidence
Keyword stuffingNull or negativeMeasurably worse than baseline across multiple benchmarks
Formatting alone, or fixed recipesPoor generalisationOccasional local gains; requires matched, multi-engine testing

The row that matters most is the second one. Position in the context window is one of the strongest effects in the literature, and it is entirely conditional on the document already having been retrieved. Every hour spent on it is an hour spent on a step that only exists if a prior step has already succeeded.

The last two rows are the ones the industry does not quote. Fixed recipes generalise poorly, which is the paper's assessment of the GEO checklist as a product. And authoritative tone, the register that half the category's content advice pushes towards, grades weak and unstable, with an explicit caution not to confuse sounding credible with being credible.

Optimising the page can make it harder to find

The most uncomfortable result in the survey comes from a 2026 study that did what almost no earlier work had done: it modelled the full pipeline, including retrieval and reranking, rather than starting from a fixed candidate set.

As reported in the survey, that study covered 171,003 documents and 2,700 queries. Optimising only the body copy of a document reduced its average presence in the top 20 by about 9%, reduced its top 10 presence after reranking by 16%, and reduced final citation by 6%.

The techniques did not fail to help. They actively hurt. Writing that reads as optimised to a retrieval system looks less like the plain, relevant prose that retrieval favours, and the document loses more at the retrieval stage than it gains at the citation stage.

A second study reported in the survey points the same way. Across roughly 1,900 queries and 16,360 documents, only three of 54 method and domain combinations were significantly positive, and none of them were in question answering. Question answering is, of course, what an answer engine does.

Both of these figures come from the survey's summary of the underlying papers rather than from reading those papers directly, and should be weighted accordingly. But the direction is consistent, and it is the opposite of the direction the category sells.

Almost nothing reads your llms.txt

The llms.txt file, a plain-text summary of a site placed at its root for AI systems to read, has become a standard recommendation. Ahrefs tested whether anything reads it, using server log data from 137,210 domains in May 2026.

Of those domains, 28% published an llms.txt file. Of the files published, 97% received zero requests during the month. Among the small share that received any traffic at all, only 19.5% of requests came from named AI tools, and bots specifically identified as AI retrieval agents accounted for around 1% of requests.

The most telling line in the study is the quietest. No requests arrived for llms.txt files that did not exist. Nothing goes looking. The file is read only when something has already been directed to the site by other means, which makes it a convenience for systems that have already found you rather than a mechanism for being found.

Ahrefs note that their sample skews technical and SEO-aware, so 28% adoption should be treated as an upper bound on the wider web. The zero-request finding is unaffected by that skew.

Ranking in the top ten no longer buys the citation

The assumption underneath a great deal of GEO advice is that AI answers draw from the top of the conventional search results, so conventional ranking remains the lever. That relationship has weakened sharply.

Ahrefs analysed the result pages for 863,000 keywords and 4 million URLs cited in Google AI Overviews, in research published in March 2026. Only 37.9% of cited URLs also appeared in the conventional top ten. In July 2025 that overlap had been around 76%.

The driver is query fan-out. Rather than answering the question as asked, the engine decomposes it into sub-queries, runs them separately, and preferentially cites pages that recur across several of them. A page that ranks eleventh for the headline query but appears across four sub-queries beats a page that ranks first for one.

This changes what a page has to be. Being the best answer to a single question is worth less than being a reliable answer to a cluster of adjacent ones. For how to track this in practice, see how to measure brand visibility in AI search.

In professional services, the citations are not on your website

Everything above concerns what happens to a page. The most consequential finding concerns whether the page matters at all.

DeltaV Digital collected 21,075 AI engine responses daily between 14 April and 13 July 2026 across ChatGPT, Perplexity, Gemini, AI Overviews and AI Mode, analysing 25,337 citations across eight brands. For the business technology services brand in that set, third-party listicles earned 61% of citations. The brand's own domain earned 0.0%. The single most-cited domain for that brand was LinkedIn.

Exhibit 2: Where the citations actually came from. Citation sources for one business technology services brand, from 21,075 AI engine responses collected 14 April to 13 July 2026 across five engines. Sample: eight brands, top 30 cited URLs per brand. Vendor first-party research, not peer reviewed, and eight brands is directional rather than representative.
Source of citationShare
Third-party listicles and rankings61%
The brand's own domain0.0%
Most-cited single domainLinkedIn (736 citations)

The same dataset carries the resolution, and it is not to abandon your website. Own domains that do get retrieved convert better than any other domain type, at 1.51 citations per retrieval once a higher-education outlier is excluded, against 1.43 for editorial domains. The honest reading of that margin is that it is thin. The constraint is not that owned content lacks credibility once an engine reaches it. The constraint is reaching it.

Which returns to the survey's finding by a different route. Retrieval is the binding step. Rewriting a page you already own does not solve retrieval. Being described accurately on the surfaces that do get retrieved might.

This is the strategic case beneath the tactical work of making sure AI describes your brand correctly, and it qualifies rather than contradicts the practical routes set out in how to get your brand mentioned by ChatGPT.

What the evidence does support

Four things, and they are unglamorous.

Be relevant, in the ordinary sense. Query-document relevance grades strong and is the primary determinant of the first citation. It is also the least fashionable recommendation in the category, because it cannot be productised.

Publish extractable specifics. Quotations, statistics, prices, dates and named figures grade moderate to strong. Numbers are the unit of citation. A page that contains a checkable figure gives an engine something to credit; a page of well-composed generality gives it nothing to hold.

Get described accurately by other people. In professional services the citation surface is third-party. Directories, listicles, professional profiles and editorial coverage carry the description of a brand into the answer, and correcting a wrong description at source is worth more than another page on the owned site.

Answer clusters, not single questions. Fan-out rewards pages that recur across related sub-queries. Depth on a genuine topic beats precision on a single keyword. On the mechanism by which engines select brands in the first place, see how AI engines decide which brands to recommend.

What the evidence does not support is the recipe. Formatting alone generalises poorly. Authoritative tone is weak and unstable. Keyword stuffing is measurably worse than doing nothing. And llms.txt, on current log evidence, is read by almost nothing.

One further piece of proportion. Similarweb data reported by Search Engine Journal in July 2026 found that only 6.8% of US ChatGPT answers included a link to an external source as of May 2026, up from roughly 1% a year earlier. Search still draws around 3.3 billion average monthly unique visitors worldwide against 655 million for AI chatbots, and 461 million of ChatGPT's 494 million users also used Google in the same window. AI answers are a real and rapidly growing surface. They are not yet a replacement for the one underneath them, and a brand being told to reallocate wholesale is being oversold.

What this means for a brand in Southeast Asia

Two regional points follow, and neither is currently being made in this market.

The first is about who is doing the advising. A Vantage review of generative engine optimisation queries in August 2026, using a general web and answer index as a directional proxy rather than querying ChatGPT, Perplexity or Gemini directly, returned only agencies selling the service, all of them operating in other markets, and no sceptical source of any kind. A Southeast Asian brand evaluating GEO is therefore being briefed largely by vendors, with little regional evidence available to it. That is not a conspiracy. It is what a thin corpus looks like, and it is a reason to discount the confidence of what you are being told rather than the substance.

The second is sharper. If third-party surfaces carry the citation, then the density of those surfaces determines how hard a wrong description is to shift. In Southeast Asia the corpus is thinner than in the United States or Europe. Fewer directories, fewer regional listicles, less editorial coverage of the professional services categories. A thin corpus propagates fast and resists correction, because there are not enough competing descriptions to dilute a wrong one. A single inaccurate profile can set how an entire market's answer engines describe a firm, and it can hold that position for a long time.

The practical consequence is an inversion of the usual regional advice. In a thick corpus, publishing more is a reasonable strategy because volume eventually shifts the average. In a thin one, corroboration beats publication. Auditing and correcting how a brand is described on the handful of surfaces that actually get retrieved is higher-value work than adding pages to a site that may not be retrieved at all.

There is an obvious tension in an article making that argument on its own domain. It is worth naming rather than hiding. Owned pages that do get retrieved convert at the highest rate of any domain type, so the case is not that publishing is pointless. It is that publishing without corroboration is a bet on the step the evidence says is the constraint. Both pieces of work run together, and a consultancy that told you otherwise while publishing constantly would be worth less of your attention, not more.

Vantage is a Singapore brand consultancy specialising in brand research, strategy, and identity design for ambitious organisations across Southeast Asia, with particular depth in healthcare, finance, government, and cultural-institution branding.

Frequently asked
questions

Does generative engine optimisation actually work?
Partly. The first systematic review of the field, published in July 2026 and covering 45 studies, found that already-retrieved content can be influenced but that no reviewed technique shows a stable, cross-platform causal effect on discoverability. Relevance, extractable specifics and third-party corroboration have support behind them. Formatting recipes, authoritative tone and keyword stuffing do not. The category is real; the marketed effect sizes are not supported by its own literature.
Where does the 40% visibility increase claim come from?
From one metric in one 2024 benchmark configuration, where a visibility score rose from 19.3 to 27.2 when quotations were added to a document. The July 2026 survey that traced the figure lists it as rejected as a general claim, on the grounds that it is a relative maximum on a single metric under a specific setup. It was never a finding about brands, categories or platforms, and it should not be quoted as one.
Does llms.txt help my brand appear in AI answers?
On current evidence, almost certainly not. Ahrefs examined server logs from 137,210 domains in May 2026 and found that 97% of published llms.txt files received zero requests that month. Of the small share that saw any traffic, only 19.5% of requests came from named AI tools. Nothing requested llms.txt files that did not exist, which means the file is read only by systems that have already arrived by other means. It costs little to publish and should not be counted on.
If my page ranks first on Google, will AI cite it?
Increasingly not. Ahrefs analysed the result pages for 863,000 keywords and 4 million URLs cited in AI Overviews and found, in research published in March 2026, that only 37.9% of cited URLs appeared in the conventional top ten. That figure was around 76% in July 2025. The cause is query fan-out: the engine breaks a question into sub-queries and favours pages that recur across them, so breadth across a topic cluster now matters more than rank on a single term.
Can optimising a page for AI make things worse?
Yes, and this is the finding most often left out. A 2026 study reported in the survey modelled the full retrieval pipeline across 171,003 documents and 2,700 queries. Optimising body copy alone reduced top-20 presence by about 9%, top-10 presence after reranking by 16%, and final citation by 6%. Content written to read as optimised appears to lose more at the retrieval stage than it gains at the citation stage.
If not my website, where should the effort go?
Towards how the brand is described on surfaces other people control. In the only citation dataset covering a professional services category, third-party listicles earned 61% of one brand's citations while its own domain earned 0.0%. Directories, professional profiles, editorial coverage and industry rankings are where the description is set. Correcting an inaccurate description at source is usually higher-value than publishing another page, particularly in a market where few competing descriptions exist to dilute a wrong one.

Wondering what to fund?

Tell us a little about your brand, and we will be in touch soon.

What brings you here?

Vantage does brand strategy and identity. We do not run media, performance marketing or campaign execution.

We will get back to you soon.

Job and survey scams. Vantage does not recruit for or operate any remote “online task” or “brand survey” work, and will never ask you to pay to earn or release commissions. Anyone offering paid task or survey work on our behalf is not us. Report suspected scams to ScamShield — call 1799 or visit scamshield.gov.sg.